aiexpert
Início / Modelos / NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
n NVIDIA Lançamento · 01 de mai. de 2026

NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

NVIDIA's 120B sparse MoE with 12B active parameters, shipped at NVFP4 4-bit precision so it inferences on a single H200 — the most aggressive cost-per-token cut from a frontier-tier open weight model this quarter.

Ficha atualizada · mai. de 2026 NVIDIA Open Model Licensemai. de 20262 benchmarks
LaboratórioNVIDIA
LicençaNVIDIA Open Model License
Lançamento mai. de 2026
Identificadornvidia/Nemotron-3-Super-120B-A12B-NVFP4
88.4
GPQA-Diamond
NVIDIA-reported, NVFP4 quantization
81.7
MMLU-Pro
self-reported

Benchmarks declarados

Números da fonte, não medição nossa

Ficha técnica

Declarado pelo fornecedor
Arquitetura
[object Object][object Object]
Licença
NVIDIA Open Model License
Identificador
nvidia/Nemotron-3-Super-120B-A12B-NVFP4
Ficha gerada por
seed

Por que importa

NVFP4 turns a Hopper-class GPU into a serious frontier-host: ~30% lower TCO than running a comparable dense 70B at BF16 in vLLM, with throughput improving in the same step. For platform teams sizing their next inference cluster, this changes the math on whether you need Blackwell B200 to host frontier weights.

The license is permissive enough for production inference but restricts derivative redistribution — read the terms before fine-tuning for a downstream product.

Para quem interessa

Platform engineers running self-hosted inference for sensitive workloads, RAG architects who hit context limits on smaller models, and compute buyers re-running 2026 capacity plans.