Soft context compression condenses a context into a few memory tokens that a frozen LLM consumes in place of the raw text, but existing compressors fix the compression ratio at training and inference: each deployed ratio requires a separately trained model, and the chosen ratio is applied uniformly to all inputs, whose...
Kai-Yan Zhao, Zhong-Tao Miao, Akiko Aizawa et al.· 1 citation
CNeo-Bench, a benchmark of 4,759 Chinese neologisms with reference definitions, is introduced, organized into five top-level categories and nine subcategories by the linguistic mechanism behind each expression, paired with a two-tier evaluation framework that separates whether a model can describe a neologism from whet...
Kai-Yan Zhao, Zhong-Tao Miao, Zhe-Yong Xie et al.· 0 citations
The role of a frozen off-the-shelf instruct model as the teacher in on-policy distillation is investigated, and a key insight is revealed: the teacher reshapes the student's policy distribution so that subsequent RL converges to a superior solution that RL alone cannot reach.
Qi Ye, Zhi-Yuan Gu, Jingjie Xia et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.