Fireworks Releases Ember-1: Matches Kimi K3's Performance with 40% Fewer Tokens
Key point
Fireworks Research releases Ember-1, a specialized model built on Kimi K3 that delivers comparable performance with 40% fewer tokens, now available as a Research Preview.
Details
Fireworks Research has introduced Ember-1, a new specialized model designed to reduce the token usage of reasoning models while maintaining high quality. Built on Kimi K3, Ember-1 achieves a 40% reduction in token consumption by optimizing reasoning traces, specifically targeting unnecessary steps while preserving critical self-reflection capabilities.
The model was evaluated on several industry benchmarks, establishing a new Pareto frontier for cost-efficiency. On Terminal Bench 2.1, Ember-1 achieved an 82.0% pass rate compared to 80.9% for Kimi K3, and on SWE-bench, it scored 92.2% versus 93.2% for K3, all while using significantly fewer tokens. The model also outperformed closed models like GPT-6 Astra and Claude Opus 5 on the Bedside Bench in terms of cost-per-task.
In live A/B tests with two customers, Ember-1 demonstrated a 39% reduction in total token usage with comparable or improved task completion rates. The model is currently available as a Research Preview on Fireworks Serverless, with training support announced for enterprises looking to customize the model for specific workloads.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.