Anthropic’s newest flagship, released 22 September 2026. Claude is the family MCP was built around, so Opus 5.5 is the baseline most servers are measured against. Run it on your own tools before you compare anything else.
Paste a server URL, pick a model, and watch it call your tools in a real conversation. No install.
No server of your own? Leave it blank and use one of the public mock servers.
MCP Evals
Can Claude Opus 5.5 actually complete tasks with your server?
One chat shows the tool calls work. An eval writes real tasks from your tool schemas, has Claude Opus 5.5 drive each one, and reports where it got the wrong answer even though every call succeeded. Pick Claude Opus 5.5 as the driver.
Claude Opus 5.5 succeeds Opus 5 as Anthropic’s flagship for demanding reasoning, coding and long-horizon agent work. Anthropic says the model matches Claude Fable 5.1 on most work and generates output more than 30% faster than Opus 5. It has a 1M-token window and accepts text, image and file input.
These are the vendor’s own claims, not measurements of ours. Run the model against your own server to find out whether they hold for your tools.
Run the error-simulation mock and ask Opus 5.5 to finish a task that hits simulate_partial_success. What matters is whether it reports the partial result honestly or retries until it can claim success. That is where long-horizon agents differ most. All the mock servers →
| Model | Model ID |
|---|---|
| Claude Opus 5.5 | anthropic/claude-opus-5.5 |