Meta releases Muse Code terminal agent and Muse Spark 1.2 coding model
Model release ● Confirmed 95% confidence first seen
Meta released Muse Code, a terminal-based AI coding agent designed to handle complex, long-horizon software engineering tasks across large codebases, powered by the new Muse Spark 1.2 model. The system uses persistent background sub-agents for parallel task handling and includes features for planning, validation, and crash recovery. Pricing is $1.25/$4.25 per million input/output tokens, or $0.10/$0.20 with data collection consent.
Decision brief
- What changed
- Meta released Muse Code, a beta terminal-based coding agent powered by its new Muse Spark 1.2 model, capable of planning, writing, validating, and parallelizing software engineering tasks across large codebases; it was benchmarked on Terminal-Bench 2.1, DeepSWE v1.1, and an internal 440-task suite.
- Why it matters
- Meta is explicitly positioning Muse Code as a cheaper alternative to OpenAI's Codex and Anthropic's Claude Code, which could pressure pricing across the enterprise coding-agent market and give CTOs/CFOs a lower-cost option for large-scale engineering workflows. However, prior-generation Muse Spark 1.1 lagged competitors on the DeepSWE leaderboard (53% vs. 73-74%), so quality-versus-cost tradeoffs remain a real decision variable, not just a cost play.
- Evidence
- Three independent outlets (MarkTechPost, TechCrunch AI, The New Stack) consistently report the release and its parallel-agent architecture; MarkTechPost and TechCrunch both confirm benchmark testing and the affordability positioning, while The New Stack provides pricing and prior-model performance context.
- What remains uncertain
- No independent benchmark results for Muse Spark 1.2 itself are cited (only prior-version 1.1 scores from July), so it's unclear whether the performance gap versus GPT-5.6 Sol and Claude Opus 5 has closed. It's also unverified how enterprise customers will weigh Meta's lower pricing against the demonstrated quality gap, or how the 800 engineer-submitted corrections feeding the separate Watermelon model relate to Muse Code's actual capabilities.
- Monitor next
- Watch for independent DeepSWE or Terminal-Bench 2.1 scores for Muse Spark 1.2 to see if Meta has closed the performance gap with GPT-5.6 Sol and Claude Opus 5.
Analytical support, not advice — assumptions and open questions stated above.