Training prompts in online reinforcement learning (RL) differ substantially in how informative they are for the current policy: some are already saturated while others are too difficult to yield reliable learning signals, yet both receive equal rollout budget under standard training. We propose an exploration-guided pr...
Yuan-Hao Yue, Qianli Ma, Cheng-Yu Wang et al.· 0 citations
Long-horizon autonomous research tasks such as machine learning engineering require systems to make interdependent decisions under a limited budget. Existing LLM-based agents typically organize candidate-solution improvement through tree, graph, or chain structures, meaning that the search process determines how inform...
Shaokang Fu, Yulong Tao, Linbo Jin et al.· 1 citation
A two-stage self-evolutionary knowledge distillation framework that equips small MLLMs with robust and adaptive tool-use behaviors and introduces weighted semantic objectives and iteratively expand competence through error-driven optimization, hybrid experience replay, and group-relative policy refinement with multi-di...
Lei Shen, Cheng-Yu Wang, Yuanjie Lyu et al.· Proceedings of the 32nd ACM...· 0 citations
Autonomous multi-modal agents are increasingly important in real-world applications due to their ability to reason about complex environments and orchestrate tool use. However, deploying multi-modal large language models (MLLMs) for tool use is often constrained by computational cost and inference latency, creating a p...
Lei Shen, Chengyu Wang, Yuanjie Lyu et al.· Proceedings of the 32nd ACM...· 0 citations
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumul...
Qi-Ming Shi, Yulong Tao, Linbo Jin et al.· arXiv.org· 2 citations
This work proposes BCSD (Bidirectional Context Self-Distillation), a framework that combines self-distillation with reinforcement learning to train LLM agents to use external skills more effectively, enabling agents to utilize external skills more effectively.
Tian Pan, Yuan Li, Hong-Da Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.