TLDRocket
Sign in

The 800 mistakes that could reshape Meta’s AI coding strategy

The New Stack Amanda Caswell Covered by 6 sources

Meta's telling engineers: fix MetaCode's bugs weekly, badges included. Over 800 fixes logged so far, but the model still lags GPT and Claude on public tests.

Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Meta has found a clever, slightly unglamorous way to train its coding AI: make its own engineers clean up after it. Maher Saba, who runs the company's Applied AI Engineering group, told staff in an internal memo to submit at least one code correction per week through MetaCode, Meta's in-house coding agent. The idea is simple enough. Every time MetaCode gets something wrong and a human engineer fixes it, Meta captures the whole exchange — the original task, the bad output, the fix, and whatever testing followed. That's a kind of data GitHub repositories don't hand over on their own, since public code shows the finished product, not the messy back-and-forth that got it there.

The numbers so far are modest but real: 7,000 weekly active users have submitted more than 800 fixes, and Meta has apparently added colored badges to internal profiles to nudge people toward contributing more. Saba says those corrections have already fed into Muse Spark 1.1 and will help post-train a future model codenamed Watermelon. What Meta hasn't explained is how heavily those fixes are weighted, or whether every patch and review automatically becomes training fodder. The memo describes something more deliberate than a blanket data grab, but the details are thin.

The pressure to make this work shows up clearly on the leaderboard. Muse Spark 1.1, launched July 9, 2026, scored 53% on the DeepSWE 1.1 benchmark, a test built around 113 long-running software engineering tasks. At launch that trailed GPT-5.5's 67% and Claude Opus 4.8's 59% — and both rivals have since pulled further ahead, with GPT-5.6 Sol hitting 73% and Claude Opus 5 reaching 74% after its July 24 release. Worth remembering: DeepSWE doesn't actually test MetaCode itself, so any internal gains from this fix-it program won't show up in that public score unless Meta says so directly.

Cost is the other half of the story. Coding agents chew through tokens fast, and multiplied across a company's worth of engineers, that adds up. Muse Spark 1.1 runs at $1.25 per million input tokens and $4.25 per million output tokens, undercutting what OpenAI and Anthropic typically charge. But that advantage got thinner fast — OpenAI cut GPT-5.6 Luna's input price by 80% to $0.20 per million tokens in late July, a move it made partly in response to cheaper Chinese open-weight competitors. So Meta's model is still the budget option, but the gap it needs to close on price is shrinking just as fast as the gap it needs to close on performance.

Meta isn't alone in betting that real production code beats synthetic benchmarks as a training signal — Alibaba recently let its Qwen model run a lengthy autonomous coding stretch with every commit posted publicly to GitHub, a different route to a similar wager. Zuckerberg has said more coding and productivity tools are coming, though whether MetaCode ever leaves the building is unclear. For now, its usefulness to Meta is measured one mistake at a time.

My take — AI-written commentary, not fact-checked reporting

Turning engineers into unpaid QA testers for the boss's chatbot, then handing out badges instead of raises, is a very on-brand way for a big tech company to solve a data problem. It might even work — real mistakes beat synthetic ones every time — but it's funny that Meta needs gamification to get 7,000 people to do the job they were already doing. Meanwhile the actual competitive math hasn't moved: cheaper doesn't help much if the frontier keeps getting cheaper too.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.