The method repurposes DR-RL trajectories, which naturally contain search histories, visited webpages, evidence snippets, and final-answer supervision, and replaces the compact snippets and webpage summaries in each trajectory with the full contents of their corresponding URLs, producing substantially longer multi-docum...
Zi-Han Wang, Hao Wang, Bo Jiang et al.· 0 citations
ADRS is introduced, a framework for constructing return-associated token-level credit for multi-turn language agents that centers and normalizes privileged token scores within each step, modulates them with a return-associated Teacher Value Advantage gate based on within-group confidence--return association, and incorp...
Ran Zhang, Guinan Chen, Chenshaodong et al.· 1 citation
The proposed LLM agent early-stopping cascade outperforms the best single-gate baseline in every model-environment pair, saving 1.5-8.8 times more compute at a 90% recall target.
Kai Ruan, Zihe Huang, Ziqi Zhou et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.