#artificial intelligence
Oct 2025
NOSA: Native and Offloadable Sparse Attention
This work proposes NOSA, a trainable sparse attention mechanism natively designed for KV cache offloading that explicitly constrains the volume of CPU-GPU KV transfers, thereby achieving low communication overhead and high decoding throughput.
Yu-Xiang Huang, Peng-Jie Wang, Ji-Cheng Han et al.
· arXiv.org · 4 citations
· ⚡1