LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training c...
Shantanu Dixit, Anson Bastos, Xu-Chao Zhang et al.· 0 citations
Pseudo Self-Distillation is presented, a framework that enables small language models to construct hierarchical memory representations by distilling behavior from a strong black-box oracle through a multi-stage training pipeline, with off-policy PSD achieving the strongest results across most conditions.
Pirzada Suhail, Meng-Lin Xia, Xu-Chao Zhang et al.· 0 citations
We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify \emph{termination-token mismatch} between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama, and Gemma, th...
Yu-Xiao Yang, Tian-Run Yu, Shang-Zhe Li et al.· 2 citations
This work introduces WebXSkill, a framework that bridges a grounding gap with executable skills, each pairing a parameterized action program with step-level natural-language guidance, and finds that better skill deployment mode depends on a model's plan and execution capability.