AWS Launches llm-d Disaggregated Inference on SageMaker HyperPod and EKS, Boosting LLM Throughput by Up to 70%
AWS launches llm-d disaggregated inference on SageMaker HyperPod and EKS, delivering up to 70% higher LLM throughput by splitting compute-heavy prefill and memory-bound decode phases across distributed GPUs using an open-source Kubernetes-native framework built on vLLM.