Meta's new Muse Spark 1.2 undercuts rivals at 20 cents—if you hand over your data
Curated by the Inblix editorial team
Meta just dropped Muse Spark 1.2, a coding-focused upgrade that doubles as the company’s first dedicated coding agent. The model improves on code generation, debugging, and reasoning across large codebases by scaling up training on programming tasks—including having its predecessor, Spark 1.1, generate its own training problems and grade the solutions. On paper, it’s a solid step up from Spark 1.1, but the benchmarks Meta chose to showcase tell a more complicated story.
The company pits Spark 1.2 against Grok 4.5, Claude Opus 5, GPT-5.6 Terra, and Gemini 3.6 Flash on Terminal-Bench 2.1, DeepSWE v1.1, and 440 internal tasks. Meta’s own methodology document concedes the test setup wasn’t optimized for competing models, which shows: Opus 5 scores roughly two percentage points higher on other leaderboards. More glaring is what’s missing. Kimi K3 appears in Meta’s methodology but is absent from the published charts—and on Terminal-Bench 2.1, K3 trails Opus 5 by a slim margin while sitting comfortably ahead of Spark 1.2. The DeepSWE comparisons also can’t be mapped to the official leaderboard because each model ran inside its own agent.
That agent is Muse Code, Meta’s answer to Claude Code and OpenAI’s Codex. It installs with one command, supports planning mode (“/plan”) and a stress-testing counterpart (“/grill”), and introduces persistent sub-agents that stick around for an entire session rather than spinning up and dying. The standout feature is crash recovery: Muse Code logs every model call, approval, and change to a local protocol file, then resumes exactly where it stopped instead of re-reading the full context. That’s a genuine workflow improvement over competitors that dump you back into a lengthy replay.
Pricing is where Meta gets aggressive. The standard tier matches Spark 1.1 at $1.25 per million input tokens and $4.25 per million output tokens. The new discount tier slashes output to 20 cents per million tokens—cheaper than most Chinese providers—if you let Meta train on your data. Western competitors charge $10 to $30 for the same privilege. It’s a fitting move for a company that makes 98% of its revenue from advertising and just watched its stock drop ten percent last week. The question isn’t whether the model is good. It’s whether developers will trade their code for a bargain.
💡 Key Takeaways
- Meta's 20-cent-per-million-token pricing undercuts nearly every Western competitor but requires users to share data for model training.
- Muse Code's persistent sub-agents and exact crash-resume logging offer a tangible workflow advantage over Claude Code and OpenAI's Codex.
- Meta's own benchmarks exclude Kimi K3 from published charts despite acknowledging it was tested, and the methodology wasn't tuned for competing models.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.