aiexpert
Home / News / Brief
Research · Aug 08, 2026, 03:04 PM · 4 sources

Meta launches Muse Spark 1.2 and Muse Code; reaches 54 on Artificial Analysis Index, matches Grok 4.5

<cite index="23-3,25-3">Meta released Muse Spark 1.2, a coding-centric update to its Muse Spark model, and Muse Code, the company's first coding agent, on August 5, 2026 as it pushes to better compete with frontier AI labs like OpenAI and Anthropic.</cite> <cite index="25-3">Meta co-trained the model with Muse Code itself, using rejection-sampled harness trajectories for goals, context compaction, and sub-agents—meaning the model was explicitly tuned to perform best inside this particular tool.</cite> <cite index="28-2,28-3">On the Artificial Analysis Intelligence Index, Muse Spark 1.2 achieved a score of 54, matching SpaceXAI's Grok 4.5 and tying for third place among U.S. companies, up from Muse Spark 1.0's score of 43 released in April and 1.1's score of 51 released in July.</cite>

<cite index="29-4">On Meta's own Terminal-Bench 2.1 coding benchmark, Muse Spark 1.2 scored 82.9%, up from 76.2% for version 1.1, with a 6.7-point improvement measured the same way on the same harness a month apart.</cite> However, <cite index="25-5">on Meta's own internal coding benchmark, Muse Spark 1.2 scored 70.6%, comfortably beating GPT-5.6 Terra at 65.4% and Gemini 3.6 Flash at 63.9%, yet still sitting nearly nine points behind Opus 5's 79.4%.</cite> <cite index="28-5">Muse Code features crash-safe local event logging that records every model call, tool execution, approval, and edit, allowing users to resume from the point of interruption if a crash occurs.</cite>

<cite index="28-3">On GDPval-AA v2, which assesses real-world task execution abilities, Muse Spark 1.2 scored 1631, with only Claude Opus 5 at 1852, Claude Fable 5 at 1743, GPT-5.6 Sol at 1730, and Kimi K3 at 1685 scoring higher.</cite> For builders evaluating coding agents, Meta's transparency in publishing second-place finishes on internal benchmarks is notable—the company positioned this as "a strong but clear second place" rather than the typical "best-in-class" framing. The model is priced at $1.25/$4.25 per million input/output tokens, but a lower-priced contributor tier at $0.10/$0.20 is available.

Sources

Everything this brief rests on
  1. 01 Primary source finance.yahoo.com
  2. 02 finance.yahoo.com finance.yahoo.com
  3. 03 venturebeat.com venturebeat.com
  4. 04 finance.biggo.com finance.biggo.com