AMD MI355X Outperforms NVIDIA B200 on Kimi K3 AI Model, Delivering 3.8× Throughput at Lower Cost
Summary
AMD's MI355X GPUs are outperforming NVIDIA's B200 on the massive 2.8 trillion parameter Kimi K3 AI model, delivering 3.8× higher throughput per node at roughly 2.4× lower cost, thanks to key engineering fixes including speculative decoding patches and attention head optimizations that cut prefill times by up to 3×.
Key Points
- Kimi K3, a massive 2.8 trillion parameter open source model, is now running on AMD MI355X GPUs at ~952 tok/s/node, delivering over 3.8× the aggregate throughput per node compared to a two-node B200 deployment at roughly 2.4× lower cost per GPU than B300s.
- Key engineering fixes unlock the performance gains, including patching a missing top-k renorm probability definition in sglang's ROCm sampling branch to enable speculative decoding, which yields ~2.2× single-stream performance improvement, and zero-padding attention heads from 12 to 16 to activate AMD's fast AITER MLA prefill kernel, cutting cold prefill time by 2–3×.
- AMD's MI355X is emerging as the clear performance-per-dollar winner for frontier-scale models like Kimi K3, where its large HBM capacity provides a measurable practical edge, and improving day-0 framework support signals that the gap with CUDA-based deployments is rapidly closing.