NVIDIA Nemotron 3.5 Lightning Lands on Amazon SageMaker JumpStart With 4x Faster Throughput for AI Agents
Summary
NVIDIA Nemotron 3.5 Lightning is now live on Amazon SageMaker JumpStart, delivering up to 4x faster throughput and 30% quicker task completion for AI agents using a hybrid 30B-parameter model that runs efficiently on a single GPU with a 1M-token context window.
Key Points
- NVIDIA Nemotron 3.5 Lightning is now available on Amazon SageMaker JumpStart, offering a high-speed open model built for high-volume agentic workloads with up to 4x higher throughput and 30% faster task completion.
- The model features a hybrid Mixture-of-Experts architecture with 30B total parameters but only 3B active per forward pass, a 1M-token context window, and DFlash speculative decoding, enabling it to run on a single supported GPU.
- Users can deploy Nemotron 3.5 Lightning directly through SageMaker Studio, the Hugging Face model page, or the SageMaker Python SDK, with support for both NVFP4 and BF16 variants across use cases including financial services, cybersecurity, telecom, and retail.