Z.ai’s first native multimodal agent model, released April 2026. It is built to look at a screen or a design, plan, and act on it, which makes it a fit for MCP servers that return screenshots.
Paste a server URL, pick a model, and watch it call your tools in a real conversation. No install.
No server of your own? Leave it blank and use one of the public mock servers.
MCP Evals
Can GLM 5V Turbo actually complete tasks with your server?
One chat shows the tool calls work. An eval writes real tasks from your tool schemas, has GLM 5V Turbo drive each one, and reports where it got the wrong answer even though every call succeeded. Pick GLM 5V Turbo as the driver.
GLM 5V Turbo pairs GLM-5-Turbo with CogViT, a vision encoder Z.ai built for it. It takes text, image and video input, has a 200K window and up to 128K tokens of output, and can think before it answers. Z.ai built it for vision-based coding and GUI agents: turning designs into front-end code, operating Android and desktop interfaces, and debugging from screenshots. Z.ai names Claude Code and OpenClaw as agents it is meant to work in.
These are the vendor’s own claims, not measurements of ours. Run the model against your own server to find out whether they hold for your tools.
Run GLM 5V Turbo and GLM 5.3 Flash on the complex-schema mock with the same prompt. Both take images, but 5V Turbo was built for them. If your server returns images, give it a task that needs one, and check whether it uses what the image shows. All the mock servers →
| Model | Model ID |
|---|---|
| GLM 5V Turbo | z-ai/glm-5v-turbo |