TLDRocket
Sign in

Mistral Vibe for Code vs Claude Code vs Cursor vs Codex: Four Agents Scored on One Scaffold-to-PR Task

MarkTechPost Asif Razzaq

Four coding agents got scored on a real scaffold-test-PR task: Mistral Vibe, Claude Code, Codex, Cursor. Mistral edged out the pack on price and control, not raw power.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Coding agents just got their most useful reality check yet: not a toy demo, but a real engineering task — build a FastAPI endpoint, write the models, generate tests, run them, fix what breaks, open a PR. That's the job description for a junior dev, and it's now the job description for four AI tools fighting over your terminal: Mistral Vibe for Code, Claude Code, OpenAI Codex, and Cursor.

The scores, out of 25, landed tighter than you'd expect. Vibe for Code came out on top at 22, with Claude Code and Codex tied at 21, and Cursor trailing at 16. But the ranking tells only part of the story. Claude Code, running Opus 4.8 with 30 lifecycle hooks and a genuinely wild proof point — Bun creator Jarred Sumner reportedly ported 750,000 lines from Zig to Rust in 11 days at a 99.8% test pass rate — is clearly the sharpest blade for raw execution. Codex matches it on capability and wins on cross-surface continuity, letting a task hop between CLI, cloud, ChatGPT app, and phone without losing its place. Neither is cheap: Anthropic's own numbers put Claude Code near $13 per developer per active day before parallel runs multiply that, and OpenAI's rate card estimates $100 to $200 a month per developer for Codex.

Vibe won on a different axis entirely: money and control. Its Pro tier runs $14.99 a month, undercutting everyone, and it's the only one of the four offering self-hosting, fine-tuning on your own code, and EU data residency, with model training opt-out by default on paid plans. The underlying model story is layered — Devstral 2 handles the CLI and IDE work at 72.2% on SWE-bench Verified, Codestral does fast completions, and heavier lifting routes to Mistral Medium 3.5 — which makes judging Vibe by any single benchmark number a mistake. Mistral also claims Devstral 2 is up to 7x more cost-efficient than Claude Sonnet, though that figure comes from Mistral itself, not an independent test.

Cursor, meanwhile, is playing a slightly different sport. It's IDE-first, tuned for fast inline editing and single-file iteration rather than long autonomous loops, and its Composer 2.5 model scored 62 on Artificial Analysis's Coding Agent Index. That's not a knock on quality so much as a mismatch of design intent — Cursor optimizes for a human staying in the loop, while the other three are built to run off on their own and come back with a pull request.

Worth flagging: none of these four scores come from actually running the prompt. MarkTechPost's methodology is built from documented features, published benchmarks, and vendor specs as of mid-July 2026 — SWE-bench Verified, SWE-Bench Pro, and Terminal-Bench are different suites entirely, and reading across them as one scale is a trap. Prices and defaults move weekly in this space, so treat the numbers as a snapshot, not gospel.

My take — AI-written commentary, not fact-checked reporting

I run TLDRocket on the assumption that open weights and self-hosting matter more long-term than a few extra benchmark points, so Vibe topping this on cost-and-control rather than pure horsepower feels like the right signal, not a fluke. Claude Code is clearly the sharper tool if you're not counting euros, but a $13-a-day habit that multiplies under parallelism is exactly the kind of vendor lock-in EU teams should be nervous about. The real tell here is that nobody actually ran the prompt — benchmark theater dressed up as a bake-off, and I'd trust it twice as much if someone just pressed go.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.