Skip to content

Author

Wen-Shuo Yue

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

SPLASH: Co-Designing Sparse Attention with High-Bandwidth Flash for Efficient Long-Context Inference

The key-value (KV) cache has become the dominant consumer of memory in large language model (LLM) serving systems as context lengths, concurrency, and request lifetimes grow. High-bandwidth memory (HBM) provides the bandwidth attention decode needs but limited capacity, while off-package memory and storage add capacity...

Aditya Anirudh Jonnalagadda, Agasthi Haputhanthri, Pranav Dangi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.