TLDRocket
Sign in

Meta releases Muse Code terminal agent and Muse Spark 1.2 coding model

Model release Confirmed 95% confidence first seen

Meta released Muse Code, a terminal-based AI coding agent designed to handle complex, long-horizon software engineering tasks across large codebases, powered by the new Muse Spark 1.2 model. The system uses persistent background sub-agents for parallel task handling and includes features for planning, validation, and crash recovery. Pricing is $1.25/$4.25 per million input/output tokens, or $0.10/$0.20 with data collection consent.

Decision brief

What changed
Meta released Muse Code, a beta terminal-based coding agent powered by its new Muse Spark 1.2 model, capable of planning, writing, validating, and parallelizing software engineering tasks across large codebases; it was benchmarked on Terminal-Bench 2.1, DeepSWE v1.1, and an internal 440-task suite.
Why it matters
Meta is explicitly positioning Muse Code as a cheaper alternative to OpenAI's Codex and Anthropic's Claude Code, which could pressure pricing across the enterprise coding-agent market and give CTOs/CFOs a lower-cost option for large-scale engineering workflows. However, prior-generation Muse Spark 1.1 lagged competitors on the DeepSWE leaderboard (53% vs. 73-74%), so quality-versus-cost tradeoffs remain a real decision variable, not just a cost play.
Affected roles
CTO CFO COO
Evidence
Three independent outlets (MarkTechPost, TechCrunch AI, The New Stack) consistently report the release and its parallel-agent architecture; MarkTechPost and TechCrunch both confirm benchmark testing and the affordability positioning, while The New Stack provides pricing and prior-model performance context.
What remains uncertain
No independent benchmark results for Muse Spark 1.2 itself are cited (only prior-version 1.1 scores from July), so it's unclear whether the performance gap versus GPT-5.6 Sol and Claude Opus 5 has closed. It's also unverified how enterprise customers will weigh Meta's lower pricing against the demonstrated quality gap, or how the 800 engineer-submitted corrections feeding the separate Watermelon model relate to Muse Code's actual capabilities.
Monitor next
Watch for independent DeepSWE or Terminal-Bench 2.1 scores for Muse Spark 1.2 to see if Meta has closed the performance gap with GPT-5.6 Sol and Claude Opus 5.

Analytical support, not advice — assumptions and open questions stated above.

Source coverage

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.