NVIDIA has launched a smaller open model designed to handle the repetitive, high-volume tasks that keep autonomous agents running, aiming to cut the cost and latency of long-running agent workloads.

Called Nemotron 3.5 Lightning, the model has 30 billion total parameters but activates only 3 billion for each token. NVIDIA says this mixture-of-experts design allows it to deliver the capacity of a larger model while using compute closer to that of a much smaller one.

Nemotron 3.5 Lightning is also designed to run locally, including on NVIDIA’s DGX Spark, Jetson systems and GeForce RTX 5090 graphics cards. It can also be deployed in data centers, giving developers options for running high-volume workloads closer to where they are generated.

The model has been trained with popular agent harnesses in mind, including OpenClaw and Hermes Agent. NVIDIA says this training helps agents make more accurate tool calls while reducing latency during repetitive tasks.#NVIDIA $NVDAB

NVDAB
NVDAB
224.41
+1.94%