Skip to content

Author

Wen-Shuai Yao

We have 6 of 8 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

NOVA-CIM: Noise- and Correlation-Tolerant Stochastic Interfaces for Analog Compute-in-Memory

Analog compute-in-memory (CIM) enables energy-efficient model acceleration, but its reliance on ADC-based readout, which directly quantizes noisy column currents, makes inference accuracy highly sensitive to analog read noise, active-row scaling, and ADC precision. In this paper, we present NOVA-CIM, a noise- and corre...

Jia-Chen Ren, Wen-Shuai Yao, Hao-Bo Liu et al. · 0 citations
Preprint Sep 2026

ASSERT: Adaptive Stochastic Sampling for Robust Diffusion Models on Analog Compute-in-Memory Hardware

This work investigates diffusion inference using a noise model calibrated and validated against measurements collected from multiple physical CIM chips, and proposes ASSERT, a training-free sampler that uses higher stochasticity early and smoothly transitions to deterministic denoising.

Yuan-Nuo Feng, Yizhe Chen, Wen-Shuai Yao et al. · 0 citations
Jul 2026

Selective KV Cache Protection for Noise-Resilient LLM Inference on Analog Compute-In-Memory Systems

A hierarchical token protection strategy is proposed that keeps sink tokens and a sliding recent-token window on a higher-precision digital path while processing the bulk KV cache on analog CIM, revealing that initial and recent tokens exhibit disproportionate vulnerability to hardware noise.

Yuan-Nuo Feng, Wen-Yong Zhou, Yuang Ma et al. · 0 citations
Jul 2026

Recall Before You Rank: Similarity-Guided Top-K Reuse for Efficient Long-Context Attention

ReTopK is a training-free method that accelerates dynamic Top-$K$ attention by reusing historical retrieval decisions and retains the complete KV cache and reuses only selected indices, rather than historical scores, attention weights, or outputs.

Wen-Shuai Yao, Wen-Yong Zhou, Hanyong Shao et al. · 0 citations
Preprint Aug 2026

When Guidance Goes Off-Scale: Recalibrating Diffusion Transformers under Analog Compute-in-Memory Nonidealities

The impact of analog CIM nonidealities on DiT sampling is characterized and a retraining-free, sampler-side recalibration that adjusts only the CFG scale for a given CIM condition is proposed, showing that the optimal guidance scale increases with CIM noise.

Wen-Shuai Yao, Wen-Yong Zhou · 0 citations
#artificial intelligence Preprint Aug 2026

Approximate Speculative Decoding

Approximate Speculative Decoding (ASD) is introduced, a training-free verifier that replaces binary first-mismatch truncation with budgeted longest-prefix selection and reuses the contiguous target-greedy suffix without additional approximate decisions or target-model forward passes.

Yuan-Nuo Feng, Zegang Peng, Yu-Xin Xie et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.