The small open model from Alibaba’s Qwen 3.8 generation, released with Apache 2.0 weights in August 2026. A 4-bit build is about 17GB and runs on a laptop, and it thinks at length unless you tell it otherwise.
Paste a server URL, pick a model, and watch it call your tools in a real conversation. No install.
No server of your own? Leave it blank and use one of the public mock servers.
MCP Evals
Can Qwen 3.8 27B actually complete tasks with your server?
One chat shows the tool calls work. An eval writes real tasks from your tool schemas, has Qwen 3.8 27B drive each one, and reports where it got the wrong answer even though every call succeeded. Pick Qwen 3.8 27B as the driver.
Qwen 3.8 27B is a dense 27B-parameter vision-language model. It mixes Gated DeltaNet linear-attention layers with Gated Attention layers, three to one, which keeps a long window affordable: 262K natively and up to 1M with YaRN scaling. It takes text, image and video input. Thinking is on by default at xhigh effort, and preserved thinking is on too, so earlier reasoning stays in the history across turns.
These are the vendor’s own claims, not measurements of ours. Run the model against your own server to find out whether they hold for your tools.
Run the same four-step task against the complex-schema mock on Qwen 3.8 27B and Qwen 3.8 Max. The 27B is a fraction of the size. If it makes the same calls, it is the one to run locally beside your server while you develop. All the mock servers →
| Model | Model ID |
|---|---|
| Qwen 3.8 27B | qwen/qwen3.8-27b |