Whitepaper: RNGD Benchmarking on Backend.AI for Efficient AI Inference
Key point
RNGD demonstrated competitive inference performance with up to 44% less power consumption compared to the RTX PRO 6000.
Details
FuriosaAI and Lablup jointly published a whitepaper evaluating the performance of an AI inference stack combining the RNGD inference accelerator with the Backend.AI orchestration platform. The tests compared four RNGD units against four NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs in a bare-metal environment.
In evaluations using the Qwen3-32B(FP8) model, RNGD achieved throughput equivalent to 95% of the comparison system at a maximum concurrency of 256 requests. Simultaneously, power consumption was 30–44% lower, and throughput per watt was 1.3–1.5 times higher across all concurrency levels.
In terms of user responsiveness, RNGD maintained TTFT(Time to First Token) under 1 second up to 32 concurrent requests. Under the same conditions, the comparison system's TTFT exceeded 2.9 seconds.
The whitepaper covers not only benchmarks but also how to operate large-scale AI inference infrastructure using RNGD and Backend.AI. Backend.AI manages various accelerators, including GPUs and NPUs, on a single platform, addressing operational challenges such as:
- Air-cooling constraints
- Increasing data center power density
- Sovereign AI compliance
- Managing LLM workloads with mixed prefill and decode stages
Deployment strategies utilize Backend.AI's session execution model and the Sokovan orchestrator. This enables the deployment and management of production-scale AI inference workloads in environments containing different accelerators.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.