AI Briefing
KO

Furiosa SDK 2026.2 Improves RNGD Throughput and Deployment Speed

·2026.05.01 06:15

Key point

SDK 2026.2 raises RNGD throughput by an average of 74.9% and strengthens deployment automation.

Details

FuriosaAI has unveiled SDK 2026.2 for RNGD. Coming after RNGD entered mass production, this update focuses on boosting both performance tuning and deployment efficiency simultaneously for enterprise customers operating agentic AI, coding agents, and high-throughput LLMs at scale. It is the second RNGD SDK release this year.

Through compiler optimizations and hybrid KV cache management, the release achieved an average 74.9% throughput improvement over the previous release on Qwen3 and EXAONE 4.0. The company explained that by better allocating memory between the prefill and decode stages, it effectively doubled the serving capacity of existing hardware without increasing power consumption.

Operational complexity was also reduced.

  • Per-Model Bucket Presets in ArtifactBuilder automate prefill/decode length (bucket) tuning, eliminating manual optimization.
  • The Prefix-Aware DP Router checks a request's tokenized prefix and routes it to the RNGD replica holding the same prefix-cache, reducing redundant computation for RAG and agentic workloads.
  • Prefix Cache Hit Deferral briefly delays requests likely to match an in-flight prefix, raising cache hit rates and lowering TTFT.

RNGD is offered as a 180W TDP PCIe card and in turnkey server configurations, and the company stated it delivers up to 3.5x higher compute density than H100-based systems in standard data center environments. The company emphasized that it will continue advancing RNGD as a high-performance inference platform that runs in 10-15kW air-cooled rack environments.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.