Nvidia's New Open-Source AI Model Runs Free on a Single GPU

“Free AI should be great for hardware. Free AI should be great for chips.” That’s how Nvidia CEO Jensen Huang frames open-source AI — not as charity, as business. On August 11, Nvidia released its own open-weight model, Nemotron 3.5 Lightning, the company’s first open model release since Huang went public with his support for open-weight AI on July 24.
Runs on one GPU, free for commercial use
Nemotron 3.5 Lightning has 30 billion parameters total, but uses a mixture-of-experts architecture that only activates 3 billion of them per inference — which is why it can run on a single consumer-grade GPU, laptop or desktop. The model was distilled from Nvidia’s larger Nemotron 3 Ultra model, and it’s free for commercial use, full stop. On performance, Nvidia claims output speeds up to 4x faster than comparable open models and roughly 30% faster task completion; paired with Anthropic’s Opus 4.8 in a hybrid setup, it can cut overall task-completion costs to roughly a third of running Opus alone. Nvidia also shipped NeMo Switchyard alongside it, a routing tool that automatically picks the most appropriate and cost-effective model for each task.
Three weeks after going public with the position, Nvidia backed it up
This release didn’t come out of nowhere. On July 24, Huang made his first-ever post on X — an industry open letter urging the U.S. government not to restrict the development of open-weight AI models. The letter launched with 25 signatories and grew to more than 150 organizations within days. Its argument was blunt: open models “strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.” Releasing Nemotron 3.5 Lightning now is Nvidia turning that three-week-old public position into an actual model anyone can download and use commercially.
The real math is still about chips
Huang says the quiet part out loud: free, open AI models lower the barrier for more people to access AI technology, and that in turn “drives demand for the company’s core business of graphics processing units, or GPUs, used to train and run AI systems.” In other words, Nvidia pushing open models runs on the same logic as a shovel seller cheering on a gold rush — the more open a model is and the more people use it, the more compute gets consumed, and compute happens to be exactly what Nvidia sells.