Analog compute-in-memory (CIM) enables energy-efficient model acceleration, but its reliance on ADC-based readout, which directly quantizes noisy column currents, makes inference accuracy highly sensitive to analog read noise, active-row scaling, and ADC precision. In this paper, we present NOVA-CIM, a noise- and corre...
Jia-Chen Ren, Wen-Shuai Yao, Hao-Bo Liu et al.· 0 citations
This work investigates diffusion inference using a noise model calibrated and validated against measurements collected from multiple physical CIM chips, and proposes ASSERT, a training-free sampler that uses higher stochasticity early and smoothly transitions to deterministic denoising.
Yuan-Nuo Feng, Yizhe Chen, Wen-Shuai Yao et al.· 0 citations
A hierarchical token protection strategy is proposed that keeps sink tokens and a sliding recent-token window on a higher-precision digital path while processing the bulk KV cache on analog CIM, revealing that initial and recent tokens exhibit disproportionate vulnerability to hardware noise.
Yuan-Nuo Feng, Wen-Yong Zhou, Yuang Ma et al.· arXiv.org· 0 citations
ReTopK is a training-free method that accelerates dynamic Top-$K$ attention by reusing historical retrieval decisions and retains the complete KV cache and reuses only selected indices, rather than historical scores, attention weights, or outputs.
The impact of analog CIM nonidealities on DiT sampling is characterized and a retraining-free, sampler-side recalibration that adjusts only the CFG scale for a given CIM condition is proposed, showing that the optimal guidance scale increases with CIM noise.
Approximate Speculative Decoding (ASD) is introduced, a training-free verifier that replaces binary first-mismatch truncation with budgeted longest-prefix selection and reuses the contiguous target-greedy suffix without additional approximate decisions or target-model forward passes.
Yuan-Nuo Feng, Zegang Peng, Yu-Xin Xie et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.