Paste a server URL and watch Kimi work through your tools. Moonshot’s open-weight agent models are built for long multi-step runs, which is exactly where most MCP servers start to show cracks.
Paste a server URL, pick a model, and watch it call your tools in a real conversation. No install.
No server of your own? Leave it blank and use one of the public mock servers.
Kimi is the strongest open-weight tool caller on this list — our guide covers its standing on MCPMark and Toolathlon, benchmarks built specifically around tool invocation rather than general reasoning. That matters for MCP because a model can be excellent at prose and mediocre at deciding which of your eight tools to call. Kimi is also the one to reach for if self-hosting is on your roadmap, since open weights mean the model you test can be the model you run.
Because Kimi's weights are open, the translation happens in whatever client you run rather than behind a vendor's API — so its behaviour is more reproducible than most models here, and more dependent on your client library's conversion quality. Two Kimi deployments can disagree simply because their MCP clients convert schemas differently.
Give Kimi the complex-schema mock and let it run unprompted for several turns. Its long chains surface servers that assume one tool call per conversation far faster than a chattier model will. All the mock servers →
| Model | Model ID | Credits / run |
|---|---|---|
| Kimi K3 | moonshotai/kimi-k3 | 12 |
| Kimi K2.7 Code | moonshotai/kimi-k2.7-code | 4 |
| Kimi K2.6 | moonshotai/kimi-k2.6 | 4 |
| Kimi K2.5 | moonshotai/kimi-k2.5 | 2 |
| Kimi K2 0905 | moonshotai/kimi-k2-0905 | 2 |