Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond improvements in the distilled model's initial accuracy. Under shared RL settings, students...
Shuai Dong, Yong-Fu Zhu, Yu-Qi Xu et al.· 0 citations
Training LLM-based search agents requires high-quality search data: tasks that demand genuine multi-hop retrieval and trajectories that use search tools effectively. Existing pipelines often depend on human-written tasks, expert demonstrations, or stronger teacher models. We present SearchMaster, a self-play framework...
Wentao Tan, Qiong Cao, Jiaqi Wang et al.· 0 citations
A benchmark that evaluates whether LLM auditors can localize, attribute, and repair search-agent failures through evidence-grounded adjudication, and proposes SearchAuditor, a multi-perspective auditing framework that effectively localizes, attributes, and repairs search-agent failures through evidence-grounded adjudic...
Zhixiang Liang, Yifei Liu, Yi-Dan Huang et al.· 0 citations
This work proposes OrthoSkillVLA, a parameter-efficient framework for continual skill learning in pretrained VLA models without demonstration replay, and introduces a lightweight feature-aware MoE decoder that better preserves prior skills while acquiring new ones.
Jia-Qi Wang, Zhou Fang, Q. Shi et al.· 0 citations
FailForge is proposed, an agentic framework that converts failed rollouts into training signal, and recovers over 26% of previously failed instances at marginal additional cost, and training Qwen3.5-4B on the augmented corpus improves the SWE-bench Verified resolve rate by 6.6 points over a strong RFT baseline.
Dongyi Lv, E. Fushun, Aichen Cai et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.