Nemotron 3 Super 120B

Runs MCP tools

NVIDIA’s mid-size open model, released 11 March 2026: 120B parameters, 12B active per token. It is free on OpenRouter, which makes it the cheapest way here to see how an open agent model handles your tools.

Vendor
NVIDIA
Released
11 March 2026
Context
262K tokens
Input
Text

Paste a server URL, pick a model, and watch it call your tools in a real conversation. No install.

No server of your own? Leave it blank and use one of the public mock servers.

MCP Evals

Can Nemotron 3 Super 120B actually complete tasks with your server?

One chat shows the tool calls work. An eval writes real tasks from your tool schemas, has Nemotron 3 Super 120B drive each one, and reports where it got the wrong answer even though every call succeeded. Pick Nemotron 3 Super 120B as the driver.

What Nemotron 3 Super 120B is

Nemotron 3 Super uses the same design as the larger Nemotron 3 Ultra: interleaved Mamba-2 and mixture-of-experts layers with a few attention layers, NVIDIA’s LatentMoE routing, and multi-token prediction for faster generation. It is text-only. The weights are open under the NVIDIA Nemotron Open Model License, and NVIDIA says the FP8 build runs on two H100s or a single B200. Reasoning can be switched on or off, or run in a low-effort mode.

What NVIDIA says

  • TauBench V2 average: 61.15.
  • Terminal Bench (hard): 25.78.
  • RULER-500 at 512K tokens: 96.09%.
  • Over 50% higher token generation than leading open models.

These are the vendor’s own claims, not measurements of ours. Run the model against your own server to find out whether they hold for your tools.

How it reaches MCP

  • MCP Playground’s Agent Studio: paste your server URL above and run Nemotron 3 Super against it in the browser, with no NVIDIA key
  • NVIDIA’s hosted API at build.nvidia.com, which is OpenAI-compatible, so MCP clients that accept a custom endpoint can point at it with a base-URL change
  • The open weights, served with vLLM using the qwen3_coder tool-call parser, behind any MCP client

What to watch for

  • The free OpenRouter endpoint accepts fewer parameters than the paid one. It lists tools, tool_choice and structured_outputs, but not stop, logprobs or the penalty settings. A client that sends stop sequences gets no effect.
  • NVIDIA supports up to 1M tokens, but OpenRouter serves 262K. Check the window of the endpoint you actually call.
  • Self-hosting needs the right tool-call parser. NVIDIA’s vLLM config uses qwen3_coder; with the wrong one, tool calls come back as plain text and the MCP client never runs them.

Run the complex-schema mock on Nemotron 3 Super twice, once with reasoning on and once with it off, using the same prompt. configure_workflow takes the most structured input of the four tools, so it shows most clearly whether thinking pays for itself on your server. All the mock servers →

Available here

ModelModel ID
Nemotron 3 Super 120Bnvidia/nemotron-3-super-120b-a12b:free

Go deeper

Frequently asked questions

Is Nemotron 3 Super free?
It is free on OpenRouter, with fewer request parameters than the paid endpoint. The weights are also open, so you can host it yourself.
Nemotron 3 Super or Nemotron 3 Ultra?
Ultra has 55B active parameters to Super’s 12B and scores higher on agent benchmarks. Super is far cheaper to run. Run the same task on both against your server and see whether the gap shows up on your tools.
Does Nemotron 3 Super support tool calling?
Yes. OpenRouter lists tools and tool_choice, and NVIDIA lists tool use among its intended uses. MCP servers work through any client that turns MCP tools into function calls.
What hardware does Nemotron 3 Super need?
NVIDIA lists two H100-80GB GPUs, or a single B200 or B300 for the FP8 build. To try it against an MCP server without that, use it here.

Sources

Nemotron 3 Super for MCP: Specs & Tool Calling — Test Free | MCP Playground