Multiverse says its 438B model is fast enough for AI agents. The benchmarks tell a more complicated story.
The New Stack Amanda Caswell
Multiverse launched a 438B model for coding agents and says compression makes it fast. The catch: it’s proprietary, and the benchmarks leave some big questions open.
Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
A 438-billion-parameter model is not the usual answer when speed matters. Multiverse Computing, the Spanish company behind CompactifAI, thinks compression can change that. On Wednesday it launched Quasar 438B, its first large-scale model, aimed at coding and enterprise agents.
The pitch is simple: make a huge model cheaper and quicker to run without gutting its usefulness. Multiverse says CompactifAI can shrink models by 80% to 95% with only a small hit to accuracy. But for Quasar, it hasn’t said how much compression was applied, what model it started from, or what hardware is needed to run it.
The numbers Quasar does have are a mixed bag. Artificial Analysis gives it an Intelligence Index score of 43 and a Terminal-Bench v2.1 score of 69.3. Multiverse says that puts it ahead of Mistral Medium 3.5 and NVIDIA Nemotron 3 Ultra on the Intelligence Index, but on Terminal-Bench it still trails Claude Opus 5, which scores 89.1. Artificial Analysis also currently records Quasar at about 183 tokens per second.
That matters because Multiverse is pitching Quasar for software engineering, technical copilots, research, and workflow automation. It has a one-million-token context window, can work in English and Spanish, and is available through the Multiverse CompactifAI API. That long memory should help with big codebases and long-running tasks. But agents do not live by throughput alone. They wait for tools, carry context forward, and make repeated calls, which is where a fast model can still feel slow.
Multiverse’s July Series C, worth $570 million, was meant to help expand its compressed-model library and commercialize the technology. Quasar is the biggest proof point so far. It is also proprietary, so developers cannot inspect the weights or run it locally, which makes the company’s speed claims harder to test in real-world agent work.
My take — AI-written commentary, not fact-checked reporting
This is the kind of European AI bet that at least has an actual argument behind it: make big models practical instead of pretending everyone needs another shiny frontier demo. The problem is that compression stories always sound better before hardware, latency, and day-to-day agent messiness show up. Proprietary plus vague on the important parts is a familiar way to ask for trust, and the market has seen that movie too many times already.
Read more about this at: The New Stack