Fireworks AI Releases 'Ember-1' Based on Kimi K3: Reduces Tokens by 40% While Maintaining Quality
Key point
Fireworks AI has released 'Ember-1' as a research preview, an inference-optimized model based on Kimi K3 that reduces token usage by approximately 35–40% without quality loss in benchmarks and real-traffic tests.
Details
Fireworks Research has launched 'Ember-1' as a research preview, a 2.78T parameter MoE model with the same architecture as Kimi K3. Focused on reducing inference tokens, the model achieved a ~39% reduction in total token usage while maintaining or improving scores in Fireworks AI's internal benchmarks (such as SWE-bench Verified, DeepSWE 1.11) and customer A/B tests. In particular, inference tokens were reduced by up to 71.3%. Ember-1 is available via the Fireworks AI serverless API, and the weights are not public. Although the price is the same as Kimi K3, cost savings are expected due to token efficiency. It is currently in the research preview stage, and future general availability will be determined based on community demand.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.