TLDRocket
Sign in

Introducing: Devstral 2 and Mistral Vibe CLI.

Mistral AI

Mistral dropped Devstral 2, a new open-weight coding model, plus a terminal tool called Vibe CLI to run it. It hits 72% on SWE-bench Verified with way fewer parameters than rivals, and it's free on the API for now.

Based on reporting by Mistral AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Mistral just put out Devstral 2, and the pitch is simple: you don't need a monster model to write good code. The flagship, a 123-billion-parameter dense transformer with a 256K context window, scores 72.2% on SWE-bench Verified. That puts it near the top of the open-weight leaderboard while running on a fraction of the compute DeepSeek V3.2 or Kimi K2 need. Mistral says Devstral 2 is roughly 8 times smaller than Kimi K2 and still competitive on real coding tasks — which is the kind of efficiency claim that actually matters once you're paying for GPU time instead of reading a benchmark chart.

There's a smaller sibling too. Devstral Small 2 packs 24 billion parameters into the same 256K context window, hits 68% on the same benchmark, and is light enough to run on a single consumer GPU or even CPU-only setups. Mistral is comparing it favorably to models five times its size, and it ships under Apache 2.0, so businesses can fine-tune and deploy it on-prem without licensing headaches. Devstral 2 itself uses a modified MIT license, which is a notch more restrictive but still firmly open.

The real test came from human evaluators using Cline as the harness. Devstral 2 beat DeepSeek V3.2 with a 42.8% win rate against a 28.6% loss rate — a clear gap in Mistral's favor. But Claude Sonnet 4.5 still won more often than not, so the closed-source ceiling hasn't been cracked yet. Mistral is upfront about this, framing Devstral 2 as the best open option rather than an outright Claude killer, and leaning on cost efficiency — up to 7 times cheaper than Sonnet on real-world tasks — as the actual selling point.

Alongside the models, Mistral released Mistral Vibe CLI, an open-source terminal agent built specifically for Devstral. It scans your repo and Git status automatically, lets you reference files with @ and run shell commands with !, and claims to cut PR review cycles in half by reasoning across your whole codebase instead of just the open file. It plugs into IDEs via the Agent Communication Protocol and is already available as a Zed extension. Partners Cline and Kilo Code are already seeing real usage — Kilo Code says Devstral 2 crossed 17 billion tokens in its first 24 hours, which for a coding model nobody outside Mistral had tested yet is a pretty loud vote of confidence.

Pricing kicks in after the free API period ends: $0.40/$2.00 per million tokens for Devstral 2, and $0.10/$0.30 for the Small version. Devstral 2 wants four H100-class GPUs minimum, while Small 2 is happy on a single RTX card or NVIDIA's DGX Spark, with NIM support coming soon.

My take — AI-written commentary, not fact-checked reporting

I'll believe the 'open models are catching closed ones' narrative more once Devstral actually beats Sonnet instead of just losing by less than DeepSeek did — but shaving 5-8x off the parameter count while staying in the same benchmark neighborhood is genuinely the more interesting story here. The industry keeps chasing bigger models when the actual unlock for most developers is something that runs on one GPU and doesn't need a data center contract, and Mistral clearly gets that better than most labs right now.

Read more about this at: Mistral AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.