Magic: AI Startup Claims Frontier-Level Pretraining for a Few Million Dollars
Trending Topics Jakob Steinschaden
Magic says it trained a frontier-like model for about $500k, then scaled the same recipe to $4M. Big claim, no outside verification, and it still hasn’t shipped anything.
Based on reporting by Trending Topics, Jakob Steinschaden — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Magic has resurfaced after almost two years of silence with a research post that goes straight for one of AI’s most expensive questions: how much compute does it really take to train a competitive language model from scratch? The Vienna-founded startup says its new pretraining recipe is now more than ten times more compute-efficient than the methods behind leading open-weight base models. That is the whole pitch, really. Not bigger clusters. Better algorithms.
The company’s headline claim is striking even by AI standards. Magic says its recipe matches DeepSeek V4 Pro using 50 times less compute, which it equates to roughly half the FLOPs used for GPT-3, or about $500,000 on Nvidia GB200 systems. It then says it scaled the same approach up tenfold to around $4 million and beat every publicly available base model on perplexity evaluations. Based on its own scaling laws, Magic says training a model of similar capability with DeepSeek V4 Pro’s recipe would cost more than $100 million.
The numbers come from Magic, and that matters. The company says it used bits-per-byte loss on held-out data rather than the usual benchmark suites, because base models have not yet gone through reinforcement learning or fine-tuning and are too sensitive to prompt wording for sampling benchmarks to be reliable. Its held-out sets included its own codebase, private code repositories acquired from other startups, recent low-citation research papers, and private math problems. To check its work, Magic says it recomputed baseline figures across several inference engines and had Fireworks perform an independent check.
Its comparisons also tell you where the company wants to stand. Magic measures itself against publicly available open weights from China and the US: DeepSeek V4 Pro and V4 Flash, Moonshot’s Kimi K2, and Nvidia’s Nemotron 3 Ultra. It does not compare itself with base models from Anthropic, Google or OpenAI, because those are not public. As a proxy, it points to Kimi K3 and Meta’s Muse Spark, saying they suggest gains of 2.5x and 3.3x over Kimi K2. The blog also notes that Nemotron 3 outperforms DeepSeek V4 Pro on pretraining metrics, which suggests weaker benchmark results may be coming from post-training rather than the base model.
There is some actual model behavior in the post too. In a short reinforcement learning run on math problems, Magic says its current model reached a 72 percent pass rate. It places that next to GPT-6 Astra at 100 percent, Claude Fable 5.1 at 98 percent and Kimi K3 at 75 percent, while noting those systems have likely seen far more RL compute. Still, none of these figures have been independently verified, and Magic has not shipped a model. The post ends by saying the team looks forward to releasing “the thing.”
My take — AI-written commentary, not fact-checked reporting
Magic is doing the classic AI-startup move: turn one giant, unverifiable claim into a recruiting magnet and let the press do the rest. The real story isn’t that it says it can train cheaply; it’s that the company is still selling future tense after years of hype, funding and very little product. That’s becoming a familiar European AI pattern: strong engineering, grand promises, and the model stays backstage.
Read more about this at: Trending Topics
Related stories
Exclusive: Paying for frontier AI models buys 4-month head start at 5x the cost
Ars Technica · 2 days ago ·
46
Large Language Models: A New Moore's Law?
Hugging Face · 4 years ago ·
37
Open-source AI is just “4 months behind” closed frontier models — and 10x cheaper
The New Stack · 2 months ago ·
48