Skip to content

Selective KV Cache Protection for Noise-Resilient LLM Inference on Analog Compute-In-Memory Systems

Jul 2026 · arXiv.org · Vol abs/2607.29076 · 0 citations · 37 references
Computer Science

TL;DR

A hierarchical token protection strategy is proposed that keeps sink tokens and a sliding recent-token window on a higher-precision digital path while processing the bulk KV cache on analog CIM, revealing that initial and recent tokens exhibit disproportionate vulnerability to hardware noise.

Abstract

Analog compute-in-memory (CIM) arrays have emerged as a promising substrate for energy-efficient LLM inference, particularly for weight-stationary computations in linear layers. However, extending analog CIM to attention mechanisms introduces a fundamental challenge: KV cache operations demand repeated in-situ weight updates, and the resulting mismatch with the weight-stationary paradigm exposes dynamic computations to significant hardware noise, a critical problem that remains largely unexplored. In this paper, we present the first systematic study of dynamic attention computation on analog CIM arrays, revealing that initial and recent tokens exhibit disproportionate vulnerability to hardware noise. Motivated by this token-level insight, we propose a hierarchical token protection strategy that keeps sink tokens and a sliding recent-token window on a higher-precision digital path while processing the bulk KV cache on analog CIM. A co-designed scheduler combines analog programming, ownership transition, and bulk-MVM tile formation to bound digital overhead. Evaluations on nine LLMs show that our approach lowers average perplexity under analog noise from 33.91 to 11.95, close to the clean baseline of 11.06, while improving dynamic-KV programming-row utilization from 23.1\% to 91.2\%.

View source

Similar papers

Preprint Sep 2026

FALCON: Fault-Tolerant Magnetic Tunnel Junction-Based In-Memory Stochastic Architecture for Reliability-Critical Edge AI Applications

As modern data-centric applications such as neural inference and sensor-edge analytics expand, they increasingly encounter the von Neumann memory wall, suffering from excessive data movement overhead and stringent energy constraints. In-Memory Computing (IMC) utilizing emerging non-volatile technologies, such as Magnet...

Farzad Razi, M. Moghadam, Sercan Aygün et al. · 0 citations
Book Open access Aug 2026

MECA-CiM: A Shared-MicroExponent-aware Configurable Analog Compute-in-Memory Macro for Efficient Inference

An analog CiM accelerator based on the SMX6 format, which extends the block floating-point representation with a lightweight microexponent shared by pairs of values is presented, demonstrating that micro-exponent-aware analog CiM with configurable granularity is an effective and practical design point for energy-effici...

Wonkyung Han, Dohyun Kim, Jihoon Park et al. · 0 citations
Book Open access Sep 2026

LLM KV-cache: To Restore or To Recompute, That Is the Question

The challenges of the restoration-recomputation trade-off are investigated and its impact on inference performance when left unaddressed, and an I/O-aware KV-cache management policy is presented that dynamically navigates this trade-off.

Amirhossein Najafizadeh, Vasily Tarasov, Alex Merenstein et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.