DeepSeek V4 - Almost at the frontier, but much cheaper
Key point
DeepSeek unveiled two 1M token context V4 preview models, pushing a low-price strategy.
Details
DeepSeek has released DeepSeek-V4-Pro and DeepSeek-V4-Flash, the first preview models in the V4 series. Both models are Mixture of Experts (MoE) architectures that support 1M token context, and are distributed under the MIT license.
- V4-Pro: 1.6T total parameters, 49B active
- V4-Flash: 284B total parameters, 13B active
They also stand out in size and versatility. Pro is 865GB on Hugging Face and Flash is 160GB, and the author believes Pro appears to be the largest open weights model currently available. It's much larger than DeepSeek V3.2, and the author expected Flash to be able to run on a 128GB M5 MacBook Pro with lightweight quantization, while Pro also seemed possible if only the necessary active experts are streamed from disk.
When testing SVG generation of a pelican riding a bicycle via OpenRouter, Flash produced a fairly stable bicycle structure and pose, while Pro was generally decent but rendered the pelican somewhat awkwardly. However, the real highlight is not the demo but the pricing.
- Flash: input $0.14/M, output $0.28/M
- Pro: input $1.74/M, output $3.48/M
Flash is even cheaper than GPT-5.4 Nano, and Pro is among the cheapest of the large frontier models. The paper explains that at 1M tokens, Pro uses only 27% of the single-token FLOPs and 10% of the KV cache compared to DeepSeek V3.2, while Flash lowers these further to 10% and 7%, respectively.
In its own benchmarks, Pro shows competitive results, while the further-scaled DeepSeek-V4-Pro-Max surpasses GPT-5.2 and Gemini-3.0-Pro, though it falls slightly short of GPT-5.4 and Gemini-3.1-Pro. DeepSeek estimated the gap with the latest frontier models to be about 3-6 months, and the author is waiting for an Unsloth quantized version expected soon.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.