DeepSeek Releases V4.1-Flash: Outperforms V4-Pro with Lower Pricing
Key point
DeepSeek has released V4.1-Flash, which offers superior performance and efficiency compared to V4-Pro, along with reduced API prices.
Details
DeepSeek has launched DeepSeek-V4.1-Flash, the smallest model in its new architecture series. This model features native visual understanding capabilities and is designed to deliver higher performance and faster inference speeds.
Asymmetric Architecture and Efficiency
V4.1-Flash is an MoE (Mixture of Experts) model with a total of 552B parameters, adopting a new Causal Encoder–Decoder structure. It maximizes efficiency by using only 8B active parameters during input and 16B during output. This reduces HBM usage to 1/4 and SSD storage space to 1/8 compared to the previous generation, significantly cutting Cache-hit costs, a major component of agent expenses.
Replacement of V4-Pro and Pricing Policy
Tests by multiple organizations showed that V4.1-Flash outperforms the flagship model V4-Pro in performance, cost, speed, and total execution time. Consequently, V4-Pro will be phased out, and from September 14, 2026, all V4-Pro requests will be routed to V4.1-Flash. DeepSeek has lowered API prices to pass on cost savings from architectural optimizations to users, while maintaining an Off-peak pricing tier at 50% of the Peak rate.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.