Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajecto...
Bo-Wen Zhang, Jun-Wei He, Mao-Qi Liu et al.· 0 citations
Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated t...
Bo-wen Zhang, Junwei He, Wen Wang et al.· arXiv.org· 0 citations
Specialize-and-Merge Online Policy Distillation (SMOPD) is proposed, a two-stage training method for multi-reward optimization that outperforms GDPO across 1.5B, 3B and 7B backbones.
Wen Wang, Jia-Hua Bao, Tu Yongsiqi et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.