NVIDIA Nemotron 3.5 Lightning Lands on Amazon SageMaker JumpStart With 4x Faster Throughput for AI Agents

Aug 18, 2026
Amazon Web Services
Article image for NVIDIA Nemotron 3.5 Lightning Lands on Amazon SageMaker JumpStart With 4x Faster Throughput for AI Agents

Summary

NVIDIA Nemotron 3.5 Lightning is now live on Amazon SageMaker JumpStart, delivering up to 4x faster throughput and 30% quicker task completion for AI agents using a hybrid 30B-parameter model that runs efficiently on a single GPU with a 1M-token context window.

Key Points

  • NVIDIA Nemotron 3.5 Lightning is now available on Amazon SageMaker JumpStart, offering a high-speed open model built for high-volume agentic workloads with up to 4x higher throughput and 30% faster task completion.
  • The model features a hybrid Mixture-of-Experts architecture with 30B total parameters but only 3B active per forward pass, a 1M-token context window, and DFlash speculative decoding, enabling it to run on a single supported GPU.
  • Users can deploy Nemotron 3.5 Lightning directly through SageMaker Studio, the Hugging Face model page, or the SageMaker Python SDK, with support for both NVFP4 and BF16 variants across use cases including financial services, cybersecurity, telecom, and retail.

Tags

Read Original Article