aiexpert
Home / Models / NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
n NVIDIA Released · May 01, 2026

NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

NVIDIA's 120B sparse MoE with 12B active parameters, shipped at NVFP4 4-bit precision so it inferences on a single H200 — the most aggressive cost-per-token cut from a frontier-tier open weight model this quarter.

Profile updated · May 2026 NVIDIA Open Model LicenseMay 20262 benchmarks
LabNVIDIA
LicenseNVIDIA Open Model License
Released May 2026
Identifiernvidia/Nemotron-3-Super-120B-A12B-NVFP4
88.4
GPQA-Diamond
NVIDIA-reported, NVFP4 quantization
81.7
MMLU-Pro
self-reported

Stated benchmarks

Source numbers, not our measurement

Technical sheet

As stated by the vendor
Architecture
[object Object][object Object]
License
NVIDIA Open Model License
Identifier
nvidia/Nemotron-3-Super-120B-A12B-NVFP4
Profile generated by
seed

Why it matters

NVFP4 turns a Hopper-class GPU into a serious frontier-host: ~30% lower TCO than running a comparable dense 70B at BF16 in vLLM, with throughput improving in the same step. For platform teams sizing their next inference cluster, this changes the math on whether you need Blackwell B200 to host frontier weights.

The license is permissive enough for production inference but restricts derivative redistribution — read the terms before fine-tuning for a downstream product.

Who should care

Platform engineers running self-hosted inference for sensitive workloads, RAG architects who hit context limits on smaller models, and compute buyers re-running 2026 capacity plans.