08/18/2026
NVIDIA released Nemotron 3.5 Lightning, aimed at a real cost problem. Long running agents burn budget calling frontier models for every step, even though most of their time goes to routine work like tool calls, validating results, and handing off to sub agents. Lightning is a 30 billion parameter mixture of experts model with only 3 billion active, trained for the agent ex*****on layer rather than reasoning. NVIDIA reports up to four times faster output than similar sized models, with speculative decoding and quantization built in. It is open source with permissive licensing.
Think it. Build it. Scale it.