Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segm...
Xin-Chen Du, Zheng-Ze Zhou, Wen-Hui Zhu et al.· 0 citations
TREK (Teacher-Routed Exploration via Forward KL), a simple staged procedure that uses distillation not for imitation but for exploration support expansion, achieves high success rates early in training while unaided GRPO requires substantially more optimization steps to reach comparable levels.
LatentMT is introduced, the first systematic study of latent-reasoning LoopLMs for machine translation that adapts a small 2.6B-parameter backbone model with lightweight training and shows that hidden-representation differences shrink along the recurrent reasoning-step axis, supporting the observed saturation in perfor...
Wei-Rui Chen, Samar M. Magdy, Chiyu Zhang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.