New Open-Source Model DeepSeek-V2.5 Combines General Chat and Coding Capabilities
Key point
DeepSeek-V2.5 merges general chat and coding, released on the web and API.
Details
DeepSeek-V2.5 is a new open-source model that combines DeepSeek-V2-0628 and DeepSeek-Coder-V2-0724, offering both general chat capabilities and code processing capabilities together. Writing and instruction-following have also been improved, and it is available on both the web and the API, maintaining backward-compatible endpoints.
Access methods continue to support deepseek-coder and deepseek-chat as before. Function Calling, FIM completion, and JSON output remain unchanged, and the overall goal is a more concise and efficient user experience.
In terms of performance, DeepSeek-V2.5 outperformed both predecessors on most benchmarks. In internal Chinese evaluations, win rates against GPT-4o mini and ChatGPT-4o-latest also improved, with particularly notable gains in content creation and Q&A.
Safety has also been adjusted. Based on internal tests, the Overall Safety Score rose from 74.4% to 82.6%, and the Safety Spillover Rate dropped from 11.3% to 4.6%. The model reportedly strengthens jailbreak defenses while reducing excessive application of safety policies to normal queries.
In the coding domain, improvements were shown on HumanEval Python and LiveCodeBench (Jan 2024 - Sep 2024), and internal DS-FIM-Eval tests recorded a 5.1% improvement. However, it was noted that room for improvement remains on HumanEval Multilingual, Aider, and SWE-verified.
- Internal DS-Arena-Code evaluations also showed a significant increase in win rate against competing models.
- For some cases with lower performance, adjusting system prompt and temperature was recommended.
- The model is now open-sourced on HuggingFace: https://huggingface.co/deepseek-ai/DeepSeek-V2.5
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.