Muse Glimmer 30B

Runs MCP tools

Meta’s open 30B model, released 10 August 2026 under Apache 2.0. It is distilled from Muse Spark 1.2 and built to run agents on a single consumer GPU or a high-end Mac.

Vendor
Meta
Released
10 August 2026
Context
131K tokens
Input
Text, image

Paste a server URL, pick a model, and watch it call your tools in a real conversation. No install.

No server of your own? Leave it blank and use one of the public mock servers.

MCP Evals

Can Muse Glimmer 30B actually complete tasks with your server?

One chat shows the tool calls work. An eval writes real tasks from your tool schemas, has Muse Glimmer 30B drive each one, and reports where it got the wrong answer even though every call succeeded. Pick Muse Glimmer 30B as the driver.

What Muse Glimmer 30B is

Muse Glimmer 30B is a dense 30B-parameter model that takes text and image input, distilled from Meta’s closed Muse Spark 1.2. The weights are on Hugging Face under Apache 2.0, which allows commercial use without the restrictions on some competing open models. At full precision it needs about 55GB of memory; Meta’s 4-bit build fits in roughly 17–20GB, so a 24GB or 32GB card such as an RTX 5090 can run it. Meta built it for always-on local agents that call tools, follow multi-step plans and recover when a call fails.

What Meta says

  • MCP-Atlas: 75.5, against 62.5 for Qwen3.6-27B and 54.2 for Gemma 4 31B.
  • SWE-bench Pro: 51.2. GAIA2: 43.3. DeepSearch QA: 74.6.
  • With its speculative-decoding draft model, 3.1× faster decoding on an RTX 5090 (74.9 to 233 tokens/s in llama.cpp), 1.8× on an M5 Max and 1.5× on an M4 Max.

These are the vendor’s own claims, not measurements of ours. Run the model against your own server to find out whether they hold for your tools.

How it reaches MCP

  • MCP Playground’s Agent Studio: paste your server URL above and run Glimmer against it in the browser before you download anything
  • The open weights, served with llama.cpp or another OpenAI-compatible server, behind any MCP client on the same machine as your server
  • Hosted endpoints on OpenRouter and other providers, for a client that accepts a custom endpoint

What to watch for

  • Meta ran the rival models in its comparison without tuning them, and no independent score exists yet. MCP-Atlas is the right benchmark for this site, but 75.5 is Meta’s number.
  • The window is 131K, far smaller than Muse Spark’s 1M. A server with large tool lists or long results fills it faster than you might expect.
  • The 4-bit build is the one that fits on consumer hardware. Hosted endpoints may serve a different precision, so tool calls that work here can differ slightly on your local copy.

Run the error-simulation mock and give Glimmer a task that hits simulate_partial_success. Meta says it recovers when a tool call fails instead of stopping. Check whether it retries sensibly and reports the partial result, or gives up. All the mock servers →

Available here

ModelModel ID
Muse Glimmer 30Bmeta/muse-glimmer-30b

Go deeper

Frequently asked questions

Can Muse Glimmer 30B run locally?
Yes. Meta’s 4-bit build needs roughly 17–20GB of memory, so it fits on a 24GB or 32GB consumer GPU or a high-end Mac. Full precision needs about 55GB.
What licence is Muse Glimmer released under?
Apache 2.0, which permits commercial use, modification and redistribution.
Is Muse Glimmer good at MCP tool calling?
Meta reports 75.5 on Scale AI’s MCP-Atlas benchmark, ahead of Qwen3.6-27B (62.5) and Gemma 4 31B (54.2) in its own comparison. That is a vendor figure, so run it against your own server here before you build on it.
Muse Glimmer 30B or Muse Spark 1.2?
Spark 1.2 is the larger closed model with a 1M window. Glimmer is the open model you can run yourself, with a 131K window. If Glimmer makes the same calls on your server, you can self-host the model driving your tools.

Sources

Muse Glimmer 30B for MCP: Open Weights & Tool Calling | MCP Playground