Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through succes...
Bing-Chen Yao, Hao-Bo Xu, Hao-Kun Lin et al.· 0 citations
First, it is demonstrated that quantization is significantly more effective in preserving trustworthiness compared to pruning, and more importantly, it is demonstrated that compressing a reliable large model via quantization can produce SLMs with superior trustworthiness and adaptability compared to using small models...
Hao-Kun Lin, Kai-Jie Zhu, Hao-Bo Xu et al.· 2 citations
Results indicate that a direct, multi-attribute 3D consistency objective, when combined with high-quality correspondences, is effective for addressing the ill-posed sparse-view reconstruction problem.
Jinqian Yang, Yichen Wu, Wanhua Li et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.