Qwen 3.8 Flash

Runs MCP tools

Alibaba’s small, cheap Qwen 3.8, released 26 August 2026 as an API model and as open weights. It activates about 6B parameters per token, so it is one of the cheapest models here that still takes images and video.

Vendor
Alibaba (Qwen)
Released
27 August 2026
Context
1M tokens
Input
Text, image, video

Paste a server URL, pick a model, and watch it call your tools in a real conversation. No install.

No server of your own? Leave it blank and use one of the public mock servers.

MCP Evals

Can Qwen 3.8 Flash actually complete tasks with your server?

One chat shows the tool calls work. An eval writes real tasks from your tool schemas, has Qwen 3.8 Flash drive each one, and reports where it got the wrong answer even though every call succeeded. Pick Qwen 3.8 Flash as the driver.

What Qwen 3.8 Flash is

Qwen 3.8 Flash is a 125B-parameter multimodal mixture-of-experts model that activates about 6B parameters per token. Alibaba released it two ways on the same day: as Qwen3.8-Flash on its cloud API, and as open weights named Qwen3.8-Flash-Next on Hugging Face and ModelScope under the qwen-community-1.0 licence. The weights have a native 262K window that extends to 1M; the API model serves 1M by default. It takes text, image and video input.

What Alibaba (Qwen) says

  • SWE-bench Pro: 62.5. DeepSWE 1.1: 58.7.
  • AndroidWorld: 84.5. CoWorkBench: 73.9.
  • MathVision: 95.7.

These are the vendor’s own claims, not measurements of ours. Run the model against your own server to find out whether they hold for your tools.

How it reaches MCP

  • MCP Playground’s Agent Studio: paste your server URL above and run 3.8 Flash against it in the browser, with no Alibaba Cloud key
  • Alibaba Cloud Model Studio’s OpenAI-compatible API, so MCP clients that accept a custom endpoint can point at it with a base-URL change
  • The Flash-Next weights, served locally with an OpenAI-compatible server, behind any MCP client

What to watch for

  • The benchmark numbers are Alibaba’s own, released alongside the model and not yet reproduced. Six billion active parameters is small for a model scoring 62.5 on SWE-bench Pro, so check it on your own tools.
  • OpenRouter lists tools, tool_choice and structured_outputs but not parallel_tool_calls. Several calls in one turn may come back one at a time.
  • The open weights and the API model share an architecture but are served differently. Behaviour you see here on the API may not match a local Flash-Next build with a 262K window.

Run the same four-step task on Qwen 3.8 Flash and Qwen 3.8 Max against the complex-schema mock: create_user_profile, process_order, analyze_data, then configure_workflow. If Flash makes the same calls, you have a far cheaper model for routine tool loops. All the mock servers →

Available here

ModelModel ID
Qwen 3.8 Flashqwen/qwen3.8-flash

Go deeper

Frequently asked questions

What is the difference between Qwen3.8-Flash and Qwen3.8-Flash-Next?
Flash-Next is the open-weights release on Hugging Face and ModelScope. Qwen3.8-Flash is the production model on Alibaba’s API. They share the architecture; the API version serves a 1M window by default, while the weights are native 262K and extend to 1M.
Is Qwen 3.8 Flash open source?
The weights are public as Qwen3.8-Flash-Next under the qwen-community-1.0 licence. Read the licence before commercial use; it is not Apache 2.0.
Does Qwen 3.8 Flash support tool calling?
Yes. OpenRouter lists tools and tool_choice, and Alibaba positions it for agent workflows. MCP servers work through any client that turns MCP tools into function calls.
Qwen 3.8 Flash or GLM 5.3 Flash?
Both are small, multimodal and have windows of 1M or more. GLM 5.3 Flash accepts parallel_tool_calls on OpenRouter and Qwen 3.8 Flash does not, which matters if your client fans out calls. Run both on your server and compare.

Sources

Qwen 3.8 Flash for MCP: Specs & Tool Calling — Test Free | MCP Playground