Furiosa SDK 2025.1.0 Released
Key point
Furiosa has released SDK 2025.1.0, which optimizes LLM inference performance and supports Tool Calling.
Details
Furiosa SDK 2025.1.0 is the third SDK update, focused on optimizing LLM deployment for RNGD. This release delivers significant performance improvements along with a user-friendly API for LLM deployment.
Key updates include the following:
- LLM Latency Optimization: In large input (30K tokens) and output (1K tokens) environments, TTFT improved by up to 11.66% and TPOT improved by up to 11.45%, increasing the efficiency of high-throughput AI workloads.
- OpenAI API Tool Calling Support: Enables models to interact with external tools and functions, facilitating the building of Agentic AI applications.
- Simplified Hugging Face Model Conversion: The new
furiosa-llm buildcommand makes it easy to convert Hugging Face models into optimized model artifacts for RNGD. - Automatic Optimization of Blocked KV Cache Allocation: Reduces memory fragmentation to maximize KV cache allocation, delivering optimal performance without separate manual tuning.
In the coming months, enhanced Tensor Parallelism, Speculative Decoding (utilizing a Draft model), Embeddings API support, a torch.compile() backend, and more are planned to be added.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.