Anthropic’s flagship from July 2026, since succeeded by Opus 5.5. It is still the Opus many MCP servers were built and tested against, which makes it the baseline to check before you move a working setup to the newer model.
Paste a server URL, pick a model, and watch it call your tools in a real conversation. No install.
No server of your own? Leave it blank and use one of the public mock servers.
MCP Evals
Can Claude Opus 5 actually complete tasks with your server?
One chat shows the tool calls work. An eval writes real tasks from your tool schemas, has Claude Opus 5 drive each one, and reports where it got the wrong answer even though every call succeeded. Pick Claude Opus 5 as the driver.
Claude Opus 5 followed Opus 4.8 on 24 July 2026, at the same price, and is built for demanding reasoning, coding and long-horizon agent work. It has a 1M-token window, up to 128K tokens of output, and takes text, image and file input. Extended thinking is on by default, with a per-request effort setting. On Scale AI’s MCP-Atlas, a benchmark of 1,000 tasks across 36 real MCP servers, it scores 85.8% at xhigh effort, second only to Muse Spark 1.1 on the public leaderboard.
These are the vendor’s own claims, not measurements of ours. Run the model against your own server to find out whether they hold for your tools.
Run the same four-step task on Opus 5 and Opus 5.5 against the complex-schema mock: create_user_profile, process_order, analyze_data, then configure_workflow. If both make the same calls, the upgrade is about speed and price for your server, not about whether it works. All the mock servers →
| Model | Model ID |
|---|---|
| Claude Opus 5 | anthropic/claude-opus-5 |