Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's...
Zeng-Huang Fu, Zhao-Yang Li, Qiu-Yuan Ai et al.· 0 citations
Tree-structured reinforcement learning trains search agents by comparing alternative continuations and propagating terminal rewards to intermediate decisions. Adaptive expansion, however, creates a statistical asymmetry: an incumbent is selected using its own generation statistic, whereas fresh siblings are sampled aft...
Zeng-Huang Fu, Ning Chen, Ming-Da Jia et al.· 0 citations
The results show that evolving skill memory is not merely an inference-time plug-in: it changes policy learning and the future training distribution while retaining value as optional external memory.
Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai et al.· arXiv.org· 1 citation
CoEvoKG is introduced, a framework that turns a knowledge graph into both a source of verifiable training tasks and a persistent evidence memory for agent evolution, closing the loop between model self evolution and knowledge accumulation.
Zhaoyang Li, Zenghuang Fu, Qiuyuan Ai et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.