Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through succes...
Bing-Chen Yao, Hao-Bo Xu, Hao-Kun Lin et al.· 0 citations
Zeroth-order (ZO) optimization offers a memory-efficient alternative for LLM fine-tuning by estimating updates only from forward evaluations of perturbed parameters, without backpropagation or activation storage. However, in billion-parameter LLMs, isotropic perturbations often waste many forward evaluations on weakly...
Yue Xie, Zhi Zheng, Yun-Peng Ba et al.· 0 citations
First, it is demonstrated that quantization is significantly more effective in preserving trustworthiness compared to pruning, and more importantly, it is demonstrated that compressing a reliable large model via quantization can produce SLMs with superior trustworthiness and adaptability compared to using small models...
Hao-Kun Lin, Kai-Jie Zhu, Hao-Bo Xu et al.· 2 citations
A domain-specific debug agent is presented that addresses three core challenges in autonomous repair: mitigating knowledge scarcity through retrieved patterns and diagnostic instrumentation, ensuring integrity through anti-cheat detection and full-coverage evaluation, and controlling cost via convergence guards and bou...
Yansong Sun, Shenxi Wu, Siyuan Chen et al.· 0 citations
KOPE is presented, an experience-driven framework for hardware kernel optimization that records optimization trajectories with correctness and performance feedback in Experience Graph Memory, then uses Active Context Management and Injection to retrieve relevant experience under a fixed token budget.
Siyuan Chen, Runlin Hou, Shenxi Wu et al.· 0 citations
These findings position ES as a distinct reasoning post-training paradigm rather than a less effective, memory-efficient alternative to GRPO, and study how hyperparameter design affects the effectiveness of ES, demonstrating that ES requires a smaller population size in a larger LLM.
Yunpeng Ba, Zhi Zheng, Yue Xie et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.