NVIDIA Releases NVFP4 Quantized Model of DeepSeek-V4-Pro-0813
Key point
NVIDIA has released the NVFP4 quantized version of DeepSeek-V4-Pro-0813, supporting a MoE architecture that activates 400B parameters out of a total of 1.65T.
Details
NVIDIA has released the NVFP4 quantized model of DeepSeek-V4-Pro-0813. This model was quantized using NVIDIA's Model Optimizer and is permitted for both commercial and non-commercial use.
Architecture and Performance
The original model, DeepSeek-V4-Pro-0813, is an autoregressive language model with a Mixture-of-Experts (MoE) architecture, employing a hybrid attention mechanism that combines Compressed Sparse Attention and Heavily Compressed Attention. While the total parameter count reaches 1.65T, only 400B parameters are activated during inference to enhance efficiency. Additionally, it includes Manifold-Constrained Attention and the DSpark speculative decoding module, strengthening agent performance and inference speed.
Input and Output Characteristics
This model supports text input and can handle system prompts and multi-turn conversations. Output supports text, structured JSON, function calling, and more. In particular, inference depth can be adjusted via the reasoning effort setting, and it is optimized for high-difficulty tasks such as mathematics, software engineering, and complex problem solving. It is designed to run in NVIDIA GPU-accelerated environments, maximizing performance through CUDA cores and related libraries.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.