AI Pulse by Inblix

Nvidia's Nemotron 3.5 Lightning hits 670 tokens/sec with just 3.6B active parameters

The Decoder · Aug 11, 2026 · 3 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Nvidia's Nemotron 3.5 Lightning hits 670 tokens/sec with just 3.6B active parameters

Nvidia just dropped Nemotron 3.5 Lightning, and the numbers tell a very specific story. This is a 31.6-billion-parameter model that only activates 3.6 billion of them at any given moment, yet it matches OpenAI’s gpt-oss-120b on the Artificial Analysis Intelligence Index with a score of 24. That’s a nine-point leap over its predecessor, the Nemotron 3 Nano, achieved with the same hybrid Mamba-Transformer architecture but dramatically better execution.

Here’s the part that actually matters for anyone building agent pipelines: throughput. In pre-release tests using NVFP4 weights, Lightning screams along at nearly 670 tokens per second. That’s almost double Google’s Gemini 3.5 Flash-Lite at 386 tokens per second. A typical Intelligence Index task wraps up in about half a minute. For comparison, Qwen3.6 35B A3B needs roughly 3.5 minutes for the same work, and Gemma 4 31B drags on for 5.8 minutes. Nvidia isn’t trying to win the raw intelligence crown here — Qwen3.6 35B A3B scores a 32 and Meta’s Muse Glimmer hits 35 on that same index. They’re explicitly targeting a different spot on the efficiency curve.

The real headline might be agentic performance. On the GDPval-AA v2 benchmark, Lightning posts an Elo rating of 824 — that’s a staggering 334-point gain over the previous Nemotron 3 Nano, and it beats both gpt-oss-120b (800) and the much larger Nemotron 3 Super (698). Terminal-Bench v2.1 tells a similar story, jumping from 7 percent to 24.3 percent. For context, Nvidia worked with partners like CodeRabbit and Harvey on post-training to juice domain-specific performance, which suggests they see this as a workhorse for production agent systems, not a general-purpose chatbot.

What’s genuinely interesting here is how Lightning validates a thesis Nvidia researchers have been pushing for a while: that sub-10-billion-active-parameter models can handle most agent workloads at one-tenth to one-thirtieth the cost of 70B+ behemoths. Lightning activates 3.6 billion parameters per step and still outperforms gpt-oss-120b on agentic benchmarks while running nearly twice as fast. That paper was widely discussed; this is the first shipping product that makes the argument impossible to ignore. The model ships under the permissive OpenMDW-1.1 license with both BF16 and NVFP4 weights available now, and serverless inference is live through DeepInfra, Fireworks, CoreWeave, and several others. Text-only, million-token context window. The proprietary frontier — Gemini 3.5 Flash-Lite at 37 points, GPT-5.6 Luna at 52 — still sits comfortably ahead on pure intelligence, but for anyone optimizing cost-per-agent-task, the math here is hard to argue with.

💡 Key Takeaways

  1. Nemotron 3.5 Lightning matches OpenAI's gpt-oss-120b on intelligence benchmarks (score 24) while using roughly a quarter of the parameters and only 3.6B active at inference.
  2. At nearly 670 tokens per second, it's almost twice as fast as Google's Gemini 3.5 Flash-Lite — the fastest measured throughput in its class.
  3. Agentic benchmarks saw the biggest leap: an 824 Elo on GDPval-AA v2 beats both gpt-oss-120b and Nvidia's own larger Nemotron 3 Super, suggesting this was purpose-built for agent pipelines rather than general chat.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles