Fireworks AI has introduced Ember-1, a specialized reasoning model designed to deliver results comparable to Kimi K3 while using substantially fewer tokens. Fireworks AI announces in an official blog post that Ember-1 cuts token use by about 40 percent overall, a reduction aimed at lowering the cost of coding and agent-based AI workflows.
Reasoning models often generate extensive internal working before producing an answer. Fireworks says this can account for more than 90 percent of generated tokens. In multi-step tasks, those long reasoning traces can become especially expensive because earlier context is repeatedly sent back to the model.
According to the company, Ember-1 was trained from Kimi K3 to retain useful planning and self-correction while removing unnecessary reasoning. Fireworks says it conducted more than 50 training experiments and 200 evaluations. It says no customer data was used to train the model.
Benchmark and customer test results
Fireworks reports that Ember-1 matched or approached Kimi K3’s highest reasoning setting across several coding benchmarks, while reducing costs. On SWE-bench Verified, for example, Ember-1 reached a 92.2 percent pass rate, compared with 93.2 percent for Kimi K3 at its maximum reasoning setting. The company says Ember-1 reduced the cost of that benchmark run by 15.5 percent.
In two customer A/B tests involving production coding workloads, Fireworks says Ember-1 used 39 percent fewer total tokens and 71.3 percent fewer reasoning tokens than Kimi K3. The reported scores were nearly identical.
Ember-1 is available as a Research Preview through Fireworks Serverless. The company also plans to offer training support for organizations that want to adapt the model to their own workloads and data.
Stay up to date
AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox: