Group-based reinforcement learning (RL) has advanced large language models (LLMs) and is increasingly extending to agentic tasks, where sparse terminal rewards make step-level credit assignment essential. Existing methods assign credit from what follows an action in sampled rollouts, but do not explicitly capture its r...
Hao-Dong Zhu, Yang-Yang Ren, Chang-Bai Li et al.· 0 citations
Group-based reinforcement learning such as GRPO trains LLM agents by comparing rollouts sampled for each task, without a learned critic. In long-horizon settings, these rollouts revisit shared anchor states, offering cross-rollout evidence for step-level credit. Ideally, step-level credit should incorporate evidence be...
Yang-Yang Ren, Hao-Dong Zhu, Lin-Lin Yang et al.· 0 citations
MaPP (Marginalized Posterior-Predictive), a unified framework for data-efficient RLVR that denoises response-level advantage estimation and improves prompt selection using a shared Beta posterior, is proposed.
Yang-Yang Ren, Hao-Dong Zhu, Sheng Xu et al.· 0 citations
A Kalman-Guided Prompt Selection method (KGPS), which reformulates prompt selection as a dynamic state estimation problem rather than static difficulty prediction, and consistently improves both final accuracy and rollout efficiency over strong baselines, establishing state-of-the-art performance among online prompt se...
Hao-Dong Zhu, Yang-Yang Ren, Yanjing Li et al.· arXiv.org· 2 citations· ⚡2
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.