NVIDIA's 120B sparse MoE with 12B active parameters, shipped at NVFP4 4-bit precision so it inferences on a single H200 — the most aggressive cost-per-token cut from a frontier-tier open weight model this quarter.
NVFP4 turns a Hopper-class GPU into a serious frontier-host: ~30% lower TCO than running a comparable dense 70B at BF16 in vLLM, with throughput improving in the same step. For platform teams sizing their next inference cluster, this changes the math on whether you need Blackwell B200 to host frontier weights.
The license is permissive enough for production inference but restricts derivative redistribution — read the terms before fine-tuning for a downstream product.
Platform engineers running self-hosted inference for sensitive workloads, RAG architects who hit context limits on smaller models, and compute buyers re-running 2026 capacity plans.