FuriosaAI Unveils RNGD Server for Efficient AI Inference
Key point
FuriosaAI has unveiled the 'NXT RNGD Server,' a turnkey solution for AI inference optimized for existing data center infrastructure.
Details
FuriosaAI has introduced the NXT RNGD Server, its first branded turnkey solution built on the RNGD accelerator. This system is optimized to integrate seamlessly into existing data center environments while delivering high performance on the latest AI workloads.
With the Furiosa SDK and Furiosa LLM runtime pre-installed, applications can be run immediately after installation. Notably, it is optimized to work over standard PCIe interconnects without requiring any proprietary fabric or specialized infrastructure, delivering up to 3.5x more compute performance per rack than GPU-based systems within the same power budget.
Key technical specifications are as follows:
- Compute performance: Up to 8 RNGD accelerators installed (4 petaFLOPS FP8 per server), dual AMD EPYC processors
- Memory: 384 GB HBM3 (12 TB/s bandwidth) and 1 TB DDR5 system memory
- Power and cooling: System power of 3 kW, air-cooled
Thanks to its low power consumption of 3 kW per system, it enables AI scaling within the power and cooling constraints of most modern data centers. LG AI Research has already adopted RNGD for inference on its EXAONE model, demonstrating performance of 60 tokens per second at a 4K context on a single server (using 4 RNGD cards).
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.