FuriosaAI Begins Mass Production of RNGD for Data Center AI
Key point
FuriosaAI has started mass shipments of RNGD, marking its full-scale entry into the enterprise AI inference market.
Details
FuriosaAI has begun mass production shipments of RNGD in earnest. Through partnerships with TSMC and ASUS, the company has already delivered its first 4,000 units, and enterprise customers can adopt these immediately either as standalone PCIe cards or as turnkey servers.
CEO June Paik emphasized that there aren't many cases of a first-principles-based architecture being converted into mass-production silicon, and stressed that RNGD can run large-scale LLMs and agentic AI with far lower energy and infrastructure burden than existing solutions. The company stated that, building on this achievement, it aims to expand toward sustainable, high-performance AI computing for every enterprise.
RNGD is an accelerator for data center inference, offering 512 INT8 TFLOPS performance and high energy efficiency. Targeting the reality that most enterprise data centers are air-cooled and limited to 15kW per rack, it presents an alternative that doesn't require overhauling existing GPU-centric infrastructure.
The product is offered in two forms.
- RNGD PCIe Card: a drop-in accelerator with 180W TDP
- NXT RNGD Server: a 4U rack-mount server configuration equipped with 8 RNGD cards
The NXT server draws only 3kW of total system power, allowing 5 servers to be stacked in a standard air-cooled rack. This delivers 20 petaFLOPS (INT8) per rack, which the company describes as 3.5x higher compute density compared to H100-based systems in standard environments.
On the software side, the company offers an SDK featuring inter-chip tensor parallelism, support for Qwen 2 and Qwen 2.5, torch.compile support, use as a vLLM replacement, and OpenAI API compatibility. In addition, pre-compiled artifacts on the Hugging Face Hub support context lengths of up to 32K tokens.
The company also presented validation cases. LG AI Research confirmed that, based on the EXAONE model, RNGD showed 2.25x better performance-per-watt compared to a comparable GPU. It also stated that it ran OpenAI's GPT-OSS 120B model optimized on 2 RNGD cards, claiming that even large-scale models can be operated with less hardware.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.