Quantization-aware distillation (QAD) restores much of the short-form question-answering performance lost to sub-3-bit quantization, yet leaves mathematical and code reasoning substantially impaired. Long generations often degenerate into repetitive loops, exhausting the decoding budget without completing a solution. W...
Yuan-Teng Chen, Zhi-Lei Liu, Peisong Wang et al.· 0 citations
Low-bit quantization of linear layers is often dominated by a small number of outlier channels. Existing smoothing, rotation, and residual-based methods can mitigate this issue, but may shift the quantization bottleneck to weights or introduce costly online operations. To address these limitations, we propose PRQuant (...
Pei-Ran Wang, An-Qi Wang, Jia-Ying Zhao et al.· 0 citations
BiSCo-LLM is presented, a codebook-free binary spherical coding framework for extreme low-bit LLM weight compression and its reported storage budget includes binary codes, neural decoders, protected-channel payloads, LoRA adapters, and metadata.
Yuantian Shao, Peisong Wang, Zhilei Liu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.