AI Inference Costs Set to Surge Fivefold by 2028 as Agentic Workflows Demand More Tokens
Summary
Gartner warns that AI inference costs per agentic workflow will surge more than fivefold by 2028, driven by an 'Inference Paradox' where cheaper tokens are fueling greater AI complexity and higher overall spending, urging product leaders to build optimized multimodel strategies before costs spiral out of control.
Key Points
- Gartner predicts that AI inference costs per agentic workflow will increase more than fivefold by 2028, as more complex AI tasks require significantly more tokens than simple chatbot interactions.
- A phenomenon called the 'Inference Paradox' is emerging, where improving token unit economics are actually driving overall AI costs higher, because each new generation of AI capability demands more — and often more expensive — tokens with no reliable cost-efficient universal model in sight.
- Product leaders are urged to develop optimized multimodel ecosystems with careful inference-tiering, routing, and orchestration strategies, as defaulting to generic autonomous AI could result in unbounded costs far exceeding those of well-optimized systems.