Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through succes...
Bing-Chen Yao, Hao-Bo Xu, Hao-Kun Lin et al.· 0 citations
Continually adapting large language models requires acquiring new knowledge while preserving previously learned capabilities. Jointly adapting model parameters and task-specific soft prompts offers a promising solution, but faces two key limitations: historical prompts may become less effective as the model evolves, wh...
Rong-Guang Ye, Zhan Zhuang, Yi-Chen Wu et al.· 0 citations
First, it is demonstrated that quantization is significantly more effective in preserving trustworthiness compared to pruning, and more importantly, it is demonstrated that compressing a reliable large model via quantization can produce SLMs with superior trustworthiness and adaptability compared to using small models...
Hao-Kun Lin, Kai-Jie Zhu, Hao-Bo Xu et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.