aiexpert
Home / News / Brief
Breaking · Aug 07, 2026, 11:04 AM · 4 sources

Meta Muse Spark 1.2 + Muse Code ships: 54 on Intelligence Index, claims 82.9% Terminal-Bench, $0.69/task

Meta Superintelligence Labs released Muse Spark 1.2 and Muse Code (beta) on August 5, 2026. Muse Spark 1.2 is a reasoning model explicitly co-trained with Muse Code, Meta's terminal coding agent. Meta claims 82.9% on Terminal-Bench 2.1 and 59% on DeepSWE 1.1. Independent evaluators (Artificial Analysis, Vals) benchmark Muse Spark 1.2 at 54 on the Intelligence Index, tied with Grok 4.5 and 6”7 points behind frontier models (Claude Opus 5: 61, Claude Fable 5: 60, GPT-5.6 Sol: 59). Vals reports $0.69/test on Terminal-Bench using Terminus 2 (common harness)—the lowest cost among its top-five models. Pricing: $1.25 input / $4.25 output per million tokens; Contributor tier at $0.10/$0.20 for data-use permission.

Muse Code implements four agent types (Coordinator, Explorer, Executor, Verifier) communicating through a shared event log that survives restarts without re-deriving context. The terminal agent supports 24–30+ hour jobs (vs typical 10–15 min API call limits), persistent async background agents, and Git-worktree isolation. Meta's internal coding benchmark shows Muse Spark 1.2 at 70.6%, behind Claude Opus 5 (79.4%) but ahead of GPT-5.6 Terra (65.4%). Caveat: the 82.9% is from Meta's own harness; independent Terminal-Bench has no verified entry yet; Meta's Muse Spark 1.1 claim of 80% came in at 76.2% verified, a 3.8-point gap.

For practitioners evaluating long-horizon coding: Muse Code's event-log restart safety and parallel subagents are novel for an open model-side offering. Cost per task matters if you run high volume; Vals' $0.69 undercuts GPT-5.6 and Claude. Meta showed confidence by publishing charts where Claude Opus 5 wins—unusual for a vendor launch. Watch whether the co-training premium (model + harness) holds in other frameworks; if not, Muse Spark 1.2 alone is tier-2 on reasoning but competitive on price.

Sources

Everything this brief rests on
  1. 01 Primary source orcarouter.ai
  2. 02 OrcaRouter: Muse Spark 1.2 analysis orcarouter.ai “Muse Spark 1.2 at 54 on Intelligence Index, $0.69 per test on Vals”
  3. 03 Artificial Analysis artificialanalysis.ai “Intelligence Index 54, GDPval-AA v2 Elo 1631 #5, hallucination 28%”
  4. 04 Kingy.ai: benchmark methodology note kingy.ai “Meta claims 82.9% Terminal-Bench; verified 1.1 at 76.2% (3.8 point gap)”