Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segm...
Xin-Chen Du, Zheng-Ze Zhou, Wen-Hui Zhu et al.· 0 citations
LatentPress is introduced, which writes conversational histories and long documents into a third representation: continuous memory tokens that a frozen decoder reads directly through its input-embedding interface, with no text reconstruction at inference.
TREK (Teacher-Routed Exploration via Forward KL), a simple staged procedure that uses distillation not for imitation but for exploration support expansion, achieves high success rates early in training while unaided GRPO requires substantially more optimization steps to reach comparable levels.