Long-horizon agentic tasks demand strong reasoning and efficient execution across successive interactions with dynamic environments. A common approach decouples high-level planning from low-level execution through separate planner and actor roles. To investigate coordination failures in these tasks, we prompt both agen...
Heng-Zhuang Li, Yi-Kai Zhang, Yu Wang et al.· 0 citations
This work introduces PluginRSI, which represents a harness as a composition of atomized plugins and organizes harness evolution around these plugins, and shows that accumulating reusable mechanisms provides an effective basis for continued harness improvement.
Yao-Rui Shi, Yu-Chun Miao, Yu-Xin Chen et al.· 0 citations
This paper proposes UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data, and hopes it will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.
Tianjie Ju, Zheng Wu, Yue-Qing Sun et al.· 1 citation
CAST (Credit Assignment from Solver Teachers), which converts value changes in a game solver's state value into solver advantages and injects them into RLVR as turn-level signals and achieves the highest average zero-shot performance on ALFWorld and WebShop.
Yu Wang, Yi-Kai Zhang, Wentao Shi et al.· arXiv.org· 0 citations
This work proposes AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning that aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space.
Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao et al.· 7 citations· ⚡1
RSTG (Recovering Learning Signals via Adaptive Teacher Guidance), which applies distillation selectively and precisely where it matters most, and substantially outperforms naive GRPO+OPD.
Zhuo-Wen Han, Jinwei Xiao, Zhengxi Lu et al.· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.