Low-bit quantization of linear layers is often dominated by a small number of outlier channels. Existing smoothing, rotation, and residual-based methods can mitigate this issue, but may shift the quantization bottleneck to weights or introduce costly online operations. To address these limitations, we propose PRQuant (...
Pei-Ran Wang, An-Qi Wang, Jia-Ying Zhao et al.· 0 citations
The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path toward efficient inference, but direct MXFP4 quantization often degrades generation quality due to two numerical issues: the clipping-underflow trade-off from power-of-two scaling...
Jianlin Yu, Jing Lin, Linghui Kong et al.· arXiv.org· 0 citations
Kaleido is presented, an algorithm hardware codesign that accelerates all operations in vDiTs by exploiting channel-wise spatiotemporal correlations in latent space and skips redundant computations by reusing partial results while preserving higher generative quality than prior methods.
Wen-Xuan Miao, Haosong Liu, Weiming Hu et al.· arXiv.org· 1 citation
BiSCo-LLM is presented, a codebook-free binary spherical coding framework for extreme low-bit LLM weight compression and its reported storage budget includes binary codes, neural decoders, protected-channel payloads, LoRA adapters, and metadata.
Yuantian Shao, Peisong Wang, Zhilei Liu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.