Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens
MarkTechPost Asif Razzaq
Fireworks AI built Ember-1 by retraining Kimi K3 to use fewer tokens. It keeps accuracy while cutting roughly 40% of the bill — but only via Fireworks’ API.
Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Fireworks AI has shipped Ember-1, a specialized model built on top of Moonshot AI’s open-weight Kimi K3. The point is blunt: keep the quality, spend far less on the thinking. Fireworks says Ember-1 learns to produce shorter reasoning traces, and that this is not the same thing as simply dialing down reasoning effort at inference time.
That distinction matters because the company says reasoning models can burn through a huge share of their output on internal thought, sometimes more than 90%. In multi-turn agent workflows, that gets expensive fast. Earlier turns get replayed later, so long traces are read and billed again and again. Fireworks says customers wanted K3’s coding ability at lower cost, but lowering the effort setting gave up too much quality.
So the research team trained the model to reason more efficiently instead. Ember-1 keeps the useful parts of self-reflection — revisiting assumptions, responding to feedback — while trimming redundant loops. Fireworks says the training mix included mathematics, coding, instruction following, conversation, search, tool use, and software engineering, across both single-shot tasks and longer multi-step interactions. It ran more than 50 training experiments and over 200 evaluations, and used task plus environment feedback to guide planning and learning.
The company is also keeping a tight grip on the result. Ember-1 is available only through Fireworks’ serverless API as a Research Preview. The weights, training code, and exact training algorithms are not public, so self-hosting is off the table for now. Fireworks says all training ran on its own serverless training system and used its own data, not customer data.
On the numbers, the pitch is straightforward. Fireworks says Ember-1 matches Kimi K3’s quality with about 40% fewer tokens, and across seven benchmarks plus two customers’ production traffic it shortened reasoning by 35% to 50% without sacrificing accuracy. In one production A/B test, output tokens fell from 49.3K to 29.9K per task, reasoning tokens dropped 71.3%, total tokens dropped 39%, and the task score barely moved: 0.753 for Ember-1 versus 0.751 for K3. The catch is that pricing per token stays the same as Kimi K3 on Fireworks, so the savings come entirely from making the model talk less to itself.
My take — AI-written commentary, not fact-checked reporting
This is the right direction, and it also underlines how much modern AI waste is hidden inside “reasoning.” If a model can do the job with fewer tokens and the same price per token, that is not magic; it is product engineering finally catching up with the bill. The awkward part is that Fireworks is selling efficiency while keeping the method locked inside its own API, which is exactly the sort of move that makes open-model fans mutter into their coffee.
Read more about this at: MarkTechPost