IndexCache Cuts DeepSeek Sparse Attention Compute by 75%, Delivering Up to 1.82× Faster AI Inference on H100 GPUs
IndexCache slashes DeepSeek Sparse Attention compute by 75% by reusing token indices across transformer layers, delivering up to 1.82× faster AI inference on H100 GPUs with negligible quality loss and zero extra memory overhead, now available as a patch for SGLang and vLLM.