Skip to content

Author

Yi-Wu Yao

We have 4 of 19 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Aug 2026

PRQuant: Permutation Residual Quantization for Low-Overhead Inference

Low-bit quantization of linear layers is often dominated by a small number of outlier channels. Existing smoothing, rotation, and residual-based methods can mitigate this issue, but may shift the quantization bottleneck to weights or introduce costly online operations. To address these limitations, we propose PRQuant (...

Pei-Ran Wang, An-Qi Wang, Jia-Ying Zhao et al. · 0 citations
Jul 2026

MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention

The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path toward efficient inference, but direct MXFP4 quantization often degrades generation quality due to two numerical issues: the clipping-underflow trade-off from power-of-two scaling...

Jianlin Yu, Jing Lin, Linghui Kong et al. · 0 citations
Jul 2026

Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations

Kaleido is presented, an algorithm hardware codesign that accelerates all operations in vDiTs by exploiting channel-wise spatiotemporal correlations in latent space and skips redundant computations by reusing partial results while preserving higher generative quality than prior methods.

Wen-Xuan Miao, Haosong Liu, Weiming Hu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.