Details
- NVIDIA introduced Nemotron 3.5 Lightning, an open 30B mixture-of-experts model with 3B active parameters aimed at always-on agents handling specialized, high-volume tasks.[4]
- The company says Lightning can deliver up to 4x the output speed of similar-sized models and reached 86% accuracy on PinchBench while finishing 10,000 tasks 35% faster than Qwen3.6 35B at similar accuracy.[3][4]
- NVIDIA says the model is designed for long-running agent workflows where most time is spent on tool calls, result checks, and delegation rather than pure reasoning.[2][4]
- The model can be post-trained with NVIDIA NeMo using domain data, tools, workflows, and policies, with NVIDIA citing gains across cybersecurity, coding, legal, and energy tasks.[2][4]
- NVIDIA also released NeMo Switchyard, an open-source model-routing library that routes complex steps to frontier models and high-volume execution to Lightning.[4]
- Lightning is open and customizable, with weights, data, and recipes available now on Hugging Face; NVIDIA also says it can run from local systems such as DGX Spark to data centers.[2][4]
- Official NVIDIA materials describe the model as commercially usable and available through NVIDIA’s ecosystem, including Hugging Face and NGC.[2][7][8]
Impact
NVIDIA is positioning Nemotron 3.5 Lightning as infrastructure for the execution layer of agentic AI, not just another general-purpose chatbot model. The open-weight release may appeal to enterprises that want lower-cost, customizable models for repetitive tasks while keeping frontier models for planning and reasoning. It also strengthens NVIDIA’s software moat by tying model performance to NeMo, NIM, and its broader deployment stack, which could deepen adoption across its hardware and cloud ecosystem.