Nvidia Launches 30B AI Model and Open-Source Routing Tool That Slashes Inference Costs by Up to 74%
Summary
Nvidia launches Nemotron 3.5 Lightning, a powerful 30-billion-parameter open AI model, paired with NeMo Switchyard, an open-source routing tool that dynamically assigns tasks to the most cost-efficient model available, slashing inference costs by up to 74% while maintaining frontier-level performance.
Key Points
- Nvidia launches Nemotron 3.5 Lightning, a 30-billion-parameter open AI model, alongside NeMo Switchyard, an open-source routing library that dynamically assigns each step of an agent workflow to the most cost-efficient model available.
- Switchyard cuts benchmark costs to roughly one-third of running a frontier model alone, with real-world tests showing LangChain achieving a 74% cost reduction and Ramp cutting costs 58% while matching frontier-level performance.
- The release positions Nvidia against both standalone routing competitors like Not Diamond and RouteLLM, and a wave of competitive open-weight Chinese models, with Nvidia betting that owning both the model and routing layers under one open license is a key differentiator.