jensen huang's first-ever X post begged for open AI. three weeks later, nvidia delivered

Nemotron 3.5 Lightning is a 30-billion-parameter model with only about 3 billion active per token — Nvidia says that's why it runs up to 4x faster…

Aliteq
Ravi Malhotra · Hardware Editor

The short version

Nemotron 3.5 Lightning is a 30B-parameter mixture-of-experts model — but only about 3B parameters activate per token (the 'A3B' in its full name), which is why it's fast enough for a single RTX GPU.

The short version

Nvidia claims up to 4x higher throughput and 30% faster agentic task completion than comparably sized open models, with third-party testing measuring around 293 tokens per second.

The short version

It's built for agentic and multi-agent workloads specifically — tool use, coding, long-context — not general chat.

The short version

It ships alongside NeMo Switchyard, an open-source router that splits tasks between cheap local models and expensive frontier ones, claiming near one-third the cost of running Anthropic's Opus 4.8…

The short version

It's Nvidia's first open-weight model since Jensen Huang's July 24 X post backing a 25-company open letter on open AI.

Verdict

If you're already running local models for agentic or coding workloads on a single RTX card and 30B-class dense models have felt sluggish, Nemotron 3.5 Lightning is worth downloading this week. It's…

Aliteq

Read the full story

jensen huang's first-ever X post begged for open AI. three weeks later, nvidia delivered

Read the full story on Aliteq