NVIDIA published Nemotron 3.5 Lightning on Hugging Face, a 30 billion parameter open-weight model that activates only 3 billion parameters per token. The architecture blends Mamba-2 state-space layers, mixture-of-experts routing and a smaller set of attention layers, a hybrid NVIDIA says suits low-latency responses inside long-running agent loops rather than single-shot chat.
The release ships in an NVFP4 quantized format, NVIDIA’s 4-bit floating point scheme, alongside a full-precision BF16 checkpoint. NVIDIA reports the model runs on a single DGX Spark or H100 GPU and supports context windows up to 1 million tokens.
Nemotron 3.5 Lightning is licensed under NVIDIA’s OpenMDW Agreement, version 1.1, which permits commercial use. NVIDIA frames it as a workhorse for the smaller, cheaper calls a larger orchestrating model delegates inside a multi-step agent system, not as a standalone frontier chatbot. All benchmark figures in the model card, including SWE-bench and GPQA Diamond scores, are NVIDIA’s own internal evaluations.
NVIDIA published the Nemotron 3.5 Lightning model card on Hugging Face on August 11, 2026.