Jul 2026
Selective KV Cache Protection for Noise-Resilient LLM Inference on Analog Compute-In-Memory Systems
A hierarchical token protection strategy is proposed that keeps sink tokens and a sliding recent-token window on a higher-precision digital path while processing the bulk KV cache on analog CIM, revealing that initial and recent tokens exhibit disproportionate vulnerability to hardware noise.
Yuan-Nuo Feng, Wen-Yong Zhou, Yuang Ma et al.
· arXiv.org · 0 citations