RNGD Outpaces RTX Pro 6000 with Latest SDK
Key point
With its latest SDK, RNGD boosted the number of users per 15kW rack to up to 2x that of the RTX Pro 6000.
Details
FuriosaAI's RNGD surpassed NVIDIA RTX Pro 6000 in real-world efficiency through SDK optimization alone. The key metric wasn't peak single-user throughput, but how many users could be reliably served per rack while maintaining the 20–40 TPS/user SLO.
Based on Qwen3-32B, the SDK's March 13 tuned version significantly boosted low-batch performance compared to the February 14, 2026 version. b64 improved by 25% from about 1,200 TPS → 1,500 TPS, b32 improved by 47% from about 750 TPS → 1,100 TPS, and the number of concurrently servable users in a specific range increased 8.2x from 5.8 to 47.5.
These improvements focused on the actual service ranges of batch 64, 32, 16, 8. As a result, on a rack-power basis, RNGD served 1.8x more users at 20 TPS, 1.9x more at 30 TPS, and 2x more at 40 TPS.
RNGD also led in latency. TPOT was similar to the RTX Pro 6000, but TTFT was generally lower, nearly half in the b8–b64 range in particular. At the 30 TPS/user SLO, the RTX Pro 6000's time to first token was 2.7–4.4 seconds, while RNGD's was 1.1–2.1 seconds, less than half.
The difference in power efficiency became even more pronounced in rack-level scalability.
- 8 RNGD servers: about 3kW
- 8 RTX Pro 6000 servers: about 6.6kW
- Per 15kW rack: 5 units of RNGD vs. 2 units of RTX Pro 6000 can be installed
Because of this, the difference in card performance combines with the difference in server density, greatly widening the number of users that can be served from the same rack. According to the article, RNGD shows about 2.5x the operational efficiency per 15kW rack, substantially reducing actual data center space and TCO.
In the example internal AI assistant scenario, at 40 TPS/user, RNGD required 44 units / 9 racks, while the RTX Pro 6000 required 49 units / 25 racks. The former consumed 132kW, the latter 323.4kW, and the gap widens further when accounting for the spread of Multi-Agent systems.
The core message is clear: for real-time AI services, the number of concurrent users satisfying the SLO and service density per power matter more than peak throughput, and RNGD is rapidly raising that bar through software optimization alone.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.