Muse Spark 1.2

Runs MCP tools

Meta’s closed reasoning model for agent work, announced 5 August 2026. It is mostly a coding upgrade on Spark 1.1, the model at the top of Scale AI’s MCP-Atlas leaderboard.

Vendor
Meta
Released
6 August 2026
Context
1.05M tokens
Input
Text, image, video, PDF

Paste a server URL, pick a model, and watch it call your tools in a real conversation. No install.

No server of your own? Leave it blank and use one of the public mock servers.

MCP Evals

Can Muse Spark 1.2 actually complete tasks with your server?

One chat shows the tool calls work. An eval writes real tasks from your tool schemas, has Muse Spark 1.2 drive each one, and reports where it got the wrong answer even though every call succeeded. Pick Muse Spark 1.2 as the driver.

What Muse Spark 1.2 is

Muse Spark 1.2 is the third Muse Spark release in four months: 1.0 in April, 1.1 in July, 1.2 in August. Artificial Analysis scores it 54 on its Intelligence Index, up from 51 for 1.1 and 43 for 1.0. It has a 1M-token window and takes text, image, video and PDF input. Unlike Llama, the weights are not public. Meta’s open model in this line is Muse Glimmer 30B, which is distilled from it.

What Meta says

  • Terminal-Bench 2.1: 82.9%, against 86.7% for Claude Opus 5.
  • DeepSWE: 59.3%, against 65.0% for Claude Opus 5.
  • Better code generation, debugging and reasoning over large codebases than Spark 1.1.

These are the vendor’s own claims, not measurements of ours. Run the model against your own server to find out whether they hold for your tools.

How it reaches MCP

  • MCP Playground’s Agent Studio: paste your server URL above and run Muse Spark 1.2 against it in the browser, with no Meta account
  • Meta’s API and OpenAI-compatible routers such as OpenRouter, behind any MCP client that accepts a custom endpoint

What to watch for

  • The MCP-Atlas lead belongs to Spark 1.1 (88.1%), not 1.2. Meta pitched 1.2 as a coding upgrade, so do not assume the tool-use score carried over. Test it.
  • OpenRouter lists tools, tool_choice and structured_outputs, but neither seed nor stop. You cannot pin a seed to make runs repeatable, so run a prompt more than once before you call a result a regression.
  • Meta’s own figures put it behind Opus 5 on every coding benchmark it published. The case for it is cost per task, so measure tokens as well as whether the task finished.

Run the same four-step task on Muse Spark 1.2 and Muse Glimmer 30B against the complex-schema mock: create_user_profile, process_order, analyze_data, then configure_workflow. Glimmer is distilled from Spark, so the steps where they differ show what the open model lost. All the mock servers →

Available here

ModelModel ID
Muse Spark 1.2meta/muse-spark-1.2

Go deeper

Frequently asked questions

Is Muse Spark 1.2 open source?
No. Muse Spark is a closed model. Meta’s open-weights model in the line is Muse Glimmer 30B, released under Apache 2.0 and distilled from Spark 1.2.
How does Muse Spark do on MCP benchmarks?
Muse Spark 1.1 leads Scale AI’s public MCP-Atlas leaderboard at 88.1%, ahead of Claude Opus 5 at 85.8%. Spark 1.2 has no published entry yet, so run it on your own server.
What changed from Muse Spark 1.1?
Meta describes 1.2 as mainly a coding upgrade. Its Artificial Analysis Intelligence Index score rose from 51 to 54.
Does Muse Spark 1.2 support tool calling?
Yes. OpenRouter lists tools and tool_choice, and Meta built it for agent tasks. MCP servers work through any client that turns MCP tools into function calls.

Sources

Muse Spark 1.2 for MCP: Specs & Tool Calling — Test Free | MCP Playground