Nvidia has released Nemotron 3.5 Lightning, the first model in its new Nemotron 3.5 lineup. The model directly succeeds the Nemotron 3 Nano 30B A3B and retains its hybrid Mamba-Transformer architecture, with 31.6 billion total parameters and only 3.6 billion active at any given time. According to the independent benchmarking platform Artificial Analysis, the model scores 24 on the Intelligence Index, a nine-point increase from its predecessor (15). That places Lightning on par with OpenAI's gpt-oss-120b (24) and just behind Nvidia's own Nemotron 3 Super (26), which is about four times larger. The smartest small models in the same size class, like Qwen3.6 35B A3B (32) and Meta's new Muse Glimmer (35), still hold a clear lead.

Nvidia is targeting a different spot on the efficiency frontier with Lightning. In pre-release tests using the final NVFP4 weights, the model hits nearly 670 tokens per second, the highest measured throughput among all compared models and almost twice as fast as Google's Gemini 3.5 Flash-Lite (386 tokens/s). A task from the Intelligence Index takes about 0.5 minutes to complete, while Qwen3.6 35B A3B needs around 3.5 minutes, and Gemma 4 31B takes roughly 5.8 minutes. At 669 tokens per second, Lightning is the fastest model in the comparison. Despite generating a similar number of tokens per task as its predecessor Nemotron 3 Nano, it delivers much better results.

Nvidia ships the model under the permissive OpenMDW-1.1 license, positioning it as a high-throughput workhorse for agent-based pipelines. Artificial Analysis reports that Nvidia worked with partners like CodeRabbit and Harvey on post-training to boost performance in specific domains. Availability includes both BF16 and NVFP4 weights. The NVFP4 variant also scores 24 on the Intelligence Index with minimal quality loss compared to the higher-precision version, according to Artificial Analysis. The reasoning model handles text only and supports a context window of one million tokens. Weights are available now, and serverless inference is offered by DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe, among others.

Source: thedecoder