Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's...
Zeng-Huang Fu, Zhao-Yang Li, Qiu-Yuan Ai et al.· 0 citations
Tree-structured reinforcement learning trains search agents by comparing alternative continuations and propagating terminal rewards to intermediate decisions. Adaptive expansion, however, creates a statistical asymmetry: an incumbent is selected using its own generation statistic, whereas fresh siblings are sampled aft...
Zeng-Huang Fu, Ning Chen, Ming-Da Jia et al.· 0 citations
Self-Supervised Skill Optimization (SSO) is introduced, a comparative framework that learns a reusable skill from unlabeled task instances alone and outperforms existing GT-free prompt optimizers on both closed-ended and open-ended tasks.
Siran Peng, Cui-Yu Yang, Tianyu Fu et al.· arXiv.org· 0 citations
WebRetriever is introduced, a large-scale benchmark encompassing 800 websites and 1,550 tasks across diverse domains, including consumer, professional, and enterprise sectors, with comprehensive coverage of user intent patterns, and NavEval (Navigation Evaluation), a novel LLM-as-Judge framework that leverages rich int...
This work systematically characterize the implicit biases introduced by low-rank adaptation during alignment and establishes two theorems showing that low-rank alignment induces preferences for parameter subspaces with flat gradients and feature subspaces robust to perturbations, providing a principled explanation for...
Mingjia Shi, Shuo Wang, Xiaobo Wang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.