NVIDIA released Nemotron 3.5 Lightning on August 11, a 30-billion-parameter mixture-of-experts open-source model designed for agentic AI workloads. The model delivers up to 4x faster output speed and enables 30% faster agentic task completion compared to other models in its class, NVIDIA claims via PinchBench benchmarks. Lightning is fully customizable and runs on local hardware including RTX PCs, DGX workstations, and Jetson edge devices, giving enterprises control over data privacy and deployment.
Alongside Lightning, NVIDIA released NeMo Switchyard, an open-source routing library that intelligently directs requests to the most capable model for each task, reducing inference costs to roughly one-third of Claude Opus 4.8 alone while maintaining frontier-level accuracy. Enterprise partners including CrowdStrike, Harvey, and CodeRabbit are already customizing Lightning for domain-specific tasks.
NVIDIA's move aligns with CEO Jensen Huang's July statement that free open-source AI is good for chip sales, positioning Nemotron against Chinese competitors like DeepSeek V3 and Alibaba Qwen3.5. The model is available on Hugging Face and via NVIDIA NIM microservices, with training data and recipes published openly for transparency and reproducibility. Architects shipping agentic workflows benefit from the cost-efficiency and local-deployment flexibility, particularly for multi-model systems requiring sub-second latency.