Long-horizon coding agents receive verifiable rewards only after completing expensive sequences of tool calls. This increases inference cost, amplifies early wrong hypotheses, and can lead to sparse terminal reward and unstable training. We introduce Contextual Early Reward (CER), which predicts terminal reward through...
Ji-Han Yao, Si-Han Zeng, Shang-Bin Feng et al.· 0 citations
This paper proposes environment evolution, which incrementally increases environment difficulty off-policy and schedules the evolved environments generation by generation during training to provide continuous learning signals.
Zhi-Yuan Fan, Ting-Ting Yu, Yu-Tong Cai et al.· 2 citations
Deep Research Pretraining (DRP), an offline framework that derives predictive navigation supervision from naturally occurring evidence structures, is introduced, an offline framework that derives predictive navigation supervision from naturally occurring evidence structures.
Jiang-Nan Zhou, Zhi-Yuan Fan, Xing Wu et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.