Preprint
Jul 2026
MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models
It is shown that a layer's sensitivity depends strongly on the bitwidths of its upstream layers and that this dependence shifts the resulting preferred bit allocation, and proposed MixQuant, a technique-agnostic adaptive framework that wraps any base quantizer, outperforms adaptive and mixed-precision baselines in every setting.
Ashitabh Misra, Madhav Agrawal, Arham Jain et al.
· 0 citations