SPLASH: Co-Designing Sparse Attention with High-Bandwidth Flash for Efficient Long-Context Inference
The key-value (KV) cache has become the dominant consumer of memory in large language model (LLM) serving systems as context lengths, concurrency, and request lifetimes grow. High-bandwidth memory (HBM) provides the bandwidth attention decode needs but limited capacity, while off-package memory and storage add capacity...