NVIDIA Nemotron 3.5 Lightning
The Execution Engine for Always-On Agents
A 30B MoE with 3B active parameters which is open, customisable, up to 4 times faster in class…plus NeMo Switchyard for routing every step to the right model.
Nemotron 3.5 Lightning and Meta AI Muse Glimmer represent parallel, complementary answers to the same practical problem: how to make always-on agents fast, affordable, and controllable without wasting frontier resources on routine work.
Long-running agents do not spend most of their life thinking. They spend it doing.
Tool calls. Result validation. Subagent handoffs. git pull. Format this JSON. Check that billing line. Scan the security alert. Call the same three tools again with slightly different arguments.
If you send every one of those steps to a frontier reasoning model, you pay frontier prices and frontier latency for work that does not need frontier cognition. Always-on agents make that mistake expensive: the execution layer dominates the token budget.
NVIDIA Nemotron 3.5 Lightning is NVIDIA’s answer for that layer — and NeMo Switchyard is how you wire it into a system of models without rewriting your agent stack.
This post is a practitioner’s guide based on NVIDIA’s launch materials for Nemotron 3.5 Lightning and NeMo Switchyard.
NVIDIA’s framing is explicit: open models win when customers need control over where AI runs, how it is customised, and how it evolves…across edge devices, PCs, workstations, data centres, and the cloud.
Lightning is not trying to be the only model in the loop. It is designed to be the high-volume specialist inside a multi-model agent.
From the technical materials, Lightning is built for the execution layer of long-running agents:
High call volume
Low latency
Tool-heavy turns
Harness generalisation (OpenClaw, Hermes Agent, and other open harnesses)
Local + hybrid deployment when privacy or cost matters
Three shifts are stacking at once:
Agents are systems of models, not single-model chat apps.
Open models + open recipes let enterprises own accuracy, privacy, and cost curves.
Routing is becoming product infrastructure — as important as the models themselves.
Nemotron 3.5 Lightning is NVIDIA doubling down on the specialist slot in that stack: small enough to deploy anywhere, open enough to customize, fast enough to survive high-volume agent loops.
NeMo Switchyard is the glue that makes “the right model for this step” an engineering default rather than a custom project.
Chief Evangelist @ Kore.ai | I’m passionate about exploring the intersection of AI and language. From Language Models, AI Agents to Agentic Applications, Development Frameworks & Data-Centric Productivity Tools, I share insights and ideas on how these technologies are shaping the future.






