Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA's dual-path quantization errors, characterizing their di...
Zun-Hai Su, Yuxuan Sun, Jian-Chao Tan et al.· 0 citations
Domain-Normalized MOPD is proposed, which keeps the routing and rescales each domain's feedback by its measured spread, and improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain.
Xin Li, Hao-Zhe Jiang, Xin-Ming Gao et al.· 0 citations
Experiments on synthetic and real-world benchmarks show that FG$^2$-GDN and its variant improve associative recall and long-context understanding over GDN and KDA, with comparable computational efficiency.
Ping-Wei Sun, Yuxuan Hu, Jian-Chao Tan et al.· arXiv.org· 1 citation· ⚡1
Decay-Aware State Compression (DASC), which derives retention horizons from model weights, selects long-horizon state units, and packs them into a ragged state checkpoint layout to integrate efficiently with tensor-parallel inference engines.
Yanzhi Yu, Ping-Wei Sun, Jian-Chao Tan et al.· 2 citations
DAMP uses both quantization-error energy and decay-based persistence to identify high-risk channels during offline calibration and stores these channels at higher precision and the remainder in INT8, the first to study post-training quantization of recurrent states in GDN and KDA based language models.
Tao Zhang, Jian-Chao Tan, Ping-Wei Sun et al.· 4 citations· ⚡2
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.