Blackwell delivers Day-0 support for DeepSeek V4
Key point
NVIDIA unveiled Blackwell Day-0 support for DeepSeek V4 with performance of 3,500TPS per second.
Details
With the release of DeepSeek V4, a 1M token context and a 1.6T parameter-class model lineup were presented.
- DeepSeek-V4-Pro: total 1.6T, active parameters 49B
- DeepSeek-V4-Flash: total 284B, active parameters 13B
- Both models support a 1M token context, and according to the API documentation, the maximum output length is 384K tokens.
Model efficiency was also greatly improved. According to the article, the new version reduced single-token inference FLOPs to 27%, and also cut KV cache usage at a 1M token context down to roughly 10%.
NVIDIA rolled out Day-0 support built around Blackwell GPUs and NVFP4. According to official materials, it delivers about 3,500 TPS per GPU on GB300 or Blackwell Ultra, and this figure could go even higher with further optimization.
Key optimization points mentioned include FP4 (MXFP4) quantization, Dynamo, CUDA kernel optimization, and parallelization techniques. FP4 is used to reduce memory traffic and sampling latency during the rollout and inference stages.
Toward the end of the article, it was noted that Huawei's Ascend 950PR and 950DT are expected to support MXFP4 instructions, suggesting the possibility that DeepSeek V4 could also be adapted to the Chinese-made AI chip ecosystem.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.