Alibaba’s small, cheap Qwen 3.8, released 26 August 2026 as an API model and as open weights. It activates about 6B parameters per token, so it is one of the cheapest models here that still takes images and video.
Paste a server URL, pick a model, and watch it call your tools in a real conversation. No install.
No server of your own? Leave it blank and use one of the public mock servers.
MCP Evals
Can Qwen 3.8 Flash actually complete tasks with your server?
One chat shows the tool calls work. An eval writes real tasks from your tool schemas, has Qwen 3.8 Flash drive each one, and reports where it got the wrong answer even though every call succeeded. Pick Qwen 3.8 Flash as the driver.
Qwen 3.8 Flash is a 125B-parameter multimodal mixture-of-experts model that activates about 6B parameters per token. Alibaba released it two ways on the same day: as Qwen3.8-Flash on its cloud API, and as open weights named Qwen3.8-Flash-Next on Hugging Face and ModelScope under the qwen-community-1.0 licence. The weights have a native 262K window that extends to 1M; the API model serves 1M by default. It takes text, image and video input.
These are the vendor’s own claims, not measurements of ours. Run the model against your own server to find out whether they hold for your tools.
Run the same four-step task on Qwen 3.8 Flash and Qwen 3.8 Max against the complex-schema mock: create_user_profile, process_order, analyze_data, then configure_workflow. If Flash makes the same calls, you have a far cheaper model for routine tool loops. All the mock servers →
| Model | Model ID |
|---|---|
| Qwen 3.8 Flash | qwen/qwen3.8-flash |