Furiosa SDK 2026.1: Hybrid Batching, Prefix Caching, Native Kubernetes Support
Key point
Furiosa SDK 2026.1 strengthens batching, caching, and Kubernetes support.
Details
FuriosaAI has released SDK 2026.1, laying the groundwork to elevate RNGD from experimental stage to enterprise-grade operational stage. This update emphasizes full-stack infrastructure for RAG and agentic AI workloads, and strengthens observability with Rust-native telemetry and OpenTelemetry support.
At the serving layer, key performance improvements stand out. Hybrid Batching combines prefill and decode requests into a single batch, boosting RPS by up to 2x, and Prefix Caching reuses common prompt prefixes via a branch-compressed radix tree to reduce TTFT.
The feature scope has also expanded. Pooling model support now covers embedding, scoring, and reranking, while Advanced Quantization offers fine-grained dynamic FP8 and DeepSeek-style 2D-block weight quantization. Additionally, ARM64 support, support for the EXAONE 4.0 and Qwen3 families, and torch.compile and vLLM-compatible APIs have been included, broadening deployment options.
On the cloud-native operations side, the llm-d framework, NPU Operator, and Dynamic Resource Allocation (DRA) have been added. These enable disaggregated serving, automatic device discovery, firmware lifecycle management, and PCIe topology-aware placement on Kubernetes, simplifying the operation of large-scale NPU clusters.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.