TLDRocket
Sign in

Ember-1 vs. Kimi K3: Nearly identical results at 3.4 times the speed

The New Stack Jessica Wachtel ● Covered by 2 sources

Ember-1 matched Kimi K3 on the tests and finished 3.4x faster. It also used fewer reasoning tokens, but Kimi can be cheaper elsewhere.

Based on reporting by The New Stack, Jessica Wachtel — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Fireworks Research launched Ember-1 as a research preview on September 23, and the pitch is simple: get Kimi K3-level quality with less thinking. Ember is built on Moonshot’s open-weight Kimi K3, and Fireworks says it learned to trim “unnecessary reasoning” while keeping the useful part. That matters because reasoning tokens are billed as output, so fewer of them should mean lower latency and lower cost.

To see whether that claim held up under pressure, the tests were set up with gradually harder prompts and repeated five times each. The same prompts were sent through OpenRouter with default reasoning settings, and both models were routed to Fireworks so the speed comparison stayed clean. The checks covered logic puzzles, deployment scheduling, and probability questions, with answers verified ahead of time.

On the logic puzzles, both models went 5-for-5. Ember averaged 13,630 reasoning tokens and 3 minutes 46 seconds, while Kimi K3 averaged 16,679 reasoning tokens and 12 minutes 26 seconds. On the scheduling problem, both again solved it every time, and Ember stayed well ahead: 1 minute 29 seconds versus Kimi’s 4 minutes 46 seconds. The token gap was smaller there, but the speed gap kept showing up.

The only miss came on the probability set. Kimi nailed all five questions on every run. Ember missed once, with a tiny arithmetic error on the first question that spread to two others. Even so, it still looked strong: 4 perfect runs out of 5, 6,242 reasoning tokens on average, and 1 minute 47 seconds per run. Across all 15 runs, Ember finished with 14 perfect results, Kimi with 15.

The part that really stands out is the pace. Ember-1 finished the full test set 3.4 times faster than Kimi K3 and used 23% fewer reasoning tokens overall. On Fireworks, that translated to $2.48 for Ember versus $3.26 for Kimi. But that price story gets muddy fast, because Kimi K3 can be bought elsewhere for as little as $1 per million input tokens and $9 per million output tokens, which would have made the same runs cheaper than Ember. So the clean takeaway is not “Ember is cheaper.” It’s that Ember looks almost as accurate as Kimi K3, and much quicker.

My take — AI-written commentary, not fact-checked reporting

This is what open-weight model competition should look like: not bigger bragging rights, but less wasted thinking and less waiting. The awkward bit is that the cheapest model on paper is often the one with the most annoying asterisk attached. Speed still wins users; pricing still wins procurement. The industry keeps pretending those are the same thing, which is adorable.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.