Claude Opus 5

Runs MCP toolsReasoning

Anthropic’s flagship from July 2026, since succeeded by Opus 5.5. It is still the Opus many MCP servers were built and tested against, which makes it the baseline to check before you move a working setup to the newer model.

Vendor
Anthropic
Released
24 July 2026
Context
1M tokens
Input
Text, image, file

Paste a server URL, pick a model, and watch it call your tools in a real conversation. No install.

No server of your own? Leave it blank and use one of the public mock servers.

MCP Evals

Can Claude Opus 5 actually complete tasks with your server?

One chat shows the tool calls work. An eval writes real tasks from your tool schemas, has Claude Opus 5 drive each one, and reports where it got the wrong answer even though every call succeeded. Pick Claude Opus 5 as the driver.

What Claude Opus 5 is

Claude Opus 5 followed Opus 4.8 on 24 July 2026, at the same price, and is built for demanding reasoning, coding and long-horizon agent work. It has a 1M-token window, up to 128K tokens of output, and takes text, image and file input. Extended thinking is on by default, with a per-request effort setting. On Scale AI’s MCP-Atlas, a benchmark of 1,000 tasks across 36 real MCP servers, it scores 85.8% at xhigh effort, second only to Muse Spark 1.1 on the public leaderboard.

What Anthropic says

  • SWE-bench Verified: 96.0%. SWE-bench Pro: 79.2%.
  • OSWorld 2.0: 70.57%, up from 55.7% for Opus 4.8.
  • ARC-AGI-3: 30.2%. Frontier-Bench: 43.3%.

These are the vendor’s own claims, not measurements of ours. Run the model against your own server to find out whether they hold for your tools.

How it reaches MCP

  • Natively. Anthropic’s Messages API can connect to remote MCP servers itself, with no translation layer. On the API the model id is claude-opus-5.
  • Through Claude Code, Claude Desktop and claude.ai, all of which act as MCP clients
  • MCP Playground’s Agent Studio: paste your server URL above and run Opus 5 against it in the browser, with no Anthropic key

What to watch for

  • Thinking is on by default. Each tool-calling turn spends reasoning tokens before the call, so a lower effort setting is worth testing on simple servers where the right tool is obvious.
  • If your server passes on Opus 5, run it on Opus 5.5 before you switch. Anthropic changed the model, not the protocol, but a prompt tuned on one Opus can pick tools differently on the next.
  • OpenRouter lists structured_outputs, temperature and stop for it but not seed. Use structured output, not seed, if you need tool arguments that come back in the same shape every run.

Run the same four-step task on Opus 5 and Opus 5.5 against the complex-schema mock: create_user_profile, process_order, analyze_data, then configure_workflow. If both make the same calls, the upgrade is about speed and price for your server, not about whether it works. All the mock servers →

Available here

ModelModel ID
Claude Opus 5anthropic/claude-opus-5

Go deeper

Frequently asked questions

Claude Opus 5 or Opus 5.5?
Opus 5.5, released 22 September 2026, is the newer model and Anthropic reports it is more than 30% faster at generating output. Opus 5 is still available, so if a production agent is tuned on it, run both on your own tools before you switch.
How does Claude Opus 5 do on MCP benchmarks?
It scores 85.8% at xhigh effort on Scale AI’s MCP-Atlas, which tests tool discovery, multi-step calls and orchestration across 36 real MCP servers. That puts it second on the public leaderboard, behind Muse Spark 1.1 at 88.1%.
Does Claude Opus 5 support MCP?
Yes, natively. Anthropic’s API can connect to remote MCP servers directly, and Claude Code, Claude Desktop and claude.ai are all MCP clients. You can also run it against any server here without an Anthropic key.
What is the context window of Claude Opus 5?
1M tokens, which is both the default and the maximum, with up to 128K tokens of output.

Sources

Claude Opus 5 for MCP: Specs & Tool Calling — Test Free | MCP Playground