NVIDIA’s mid-size open model, released 11 March 2026: 120B parameters, 12B active per token. It is free on OpenRouter, which makes it the cheapest way here to see how an open agent model handles your tools.
Paste a server URL, pick a model, and watch it call your tools in a real conversation. No install.
No server of your own? Leave it blank and use one of the public mock servers.
MCP Evals
Can Nemotron 3 Super 120B actually complete tasks with your server?
One chat shows the tool calls work. An eval writes real tasks from your tool schemas, has Nemotron 3 Super 120B drive each one, and reports where it got the wrong answer even though every call succeeded. Pick Nemotron 3 Super 120B as the driver.
Nemotron 3 Super uses the same design as the larger Nemotron 3 Ultra: interleaved Mamba-2 and mixture-of-experts layers with a few attention layers, NVIDIA’s LatentMoE routing, and multi-token prediction for faster generation. It is text-only. The weights are open under the NVIDIA Nemotron Open Model License, and NVIDIA says the FP8 build runs on two H100s or a single B200. Reasoning can be switched on or off, or run in a low-effort mode.
These are the vendor’s own claims, not measurements of ours. Run the model against your own server to find out whether they hold for your tools.
Run the complex-schema mock on Nemotron 3 Super twice, once with reasoning on and once with it off, using the same prompt. configure_workflow takes the most structured input of the four tools, so it shows most clearly whether thinking pays for itself on your server. All the mock servers →
| Model | Model ID |
|---|---|
| Nemotron 3 Super 120B | nvidia/nemotron-3-super-120b-a12b:free |