Qwen 3.8 Achieves 288k Tokens/s on GB300
·2026.08.17 02:57
Key point
The Qwen 3.8 2.4T model achieved a throughput of 288,000 tokens per second in an NVIDIA GB300 NVL72 environment.
Details
NVIDIA released performance metrics for running the Qwen 3.8 2.4T model on the GB300 NVL72 system.
Key performance metrics are as follows:
- Achieved throughput of over 4,000 tokens/second per GPU (totaling 288,000 tokens per second across 72 GPUs)
- Provided throughput of over 350 tokens/second per user
- Implemented high performance using FP8 precision without separate model tuning
Performance is expected to improve further through additional optimizations, including NVFP4 precision.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.