Group Relative Return Policy Optimization (GR2PO), a critic-free reinforcement learning framework for continuous robot control, is proposed and significantly outperforms critic-free baselines that use immediate rewards and performs competitively against state-of-the-art actor-critic methods.
Abstract
Actor-critic architecture has been widely used in continuous robot control. However, they rely on learning a value network, introducing additional computational overhead during training. Moreover, policy learning may also be affected by the approximation error of value estimation. Critic-free group relative policy optimization methods provide a simpler training approach by removing the need for a critic. However, they fail to learn long-term action outcomes when directly applying immediate rewards to policy optimization in dense-reward environments. To address these problems, we propose Group Relative Return Policy Optimization (GR2PO), a critic-free reinforcement learning framework for continuous robot control. GR2PO estimates the discounted returns from the parallelly collected trajectories, performs group normalization at each rollout time index, and uses relative advantages and clipped targets to update the policy. To evaluate the effectiveness of the proposed framework, we instantiate it on robot control simulation environments and deploy the model to a real-world edge device. The results show that GR2PO significantly outperforms critic-free baselines that use immediate rewards and performs competitively against state-of-the-art actor-critic methods. Furthermore, GR2PO demonstrates competitive training efficiency. Inference tests on NVIDIA Jetson TX2 demonstrate the feasibility of deploying the learned policies on edge platforms. Further ablation experiments analyze the effects of parallel group size, return estimation methods, and target clipping ratio on learning performance. To support follow-up research, we will make the complete code publicly available after the paper is accepted, including the framework implementation, experimental configuration, and training and evaluation scripts.
A unified framework that combines centralized training with decentralized execution (CTDE) and a Hybrid Reward Architecture (HRA) is introduced that enables multiple actors to share a centralized multi-head critic and substantially improves both sample efficiency and policy performance.
Changhao Li, Yifang Zhang, Heng Zhang et al.· 0 citations
Planning with learned world models combines online trajectory optimization with learned value and policy functions for high-dimensional control. Because the planner determines the experience used for learning, while the learned critic and actor in turn score and propose future plans, planning and learning form a closed...
This work proposes Online Residual Policy Adaptation (ORPA), a framework that enables immediate, feedback-driven correction of robot actions without modifying the underlying policy parameters.
Muhammad A. Muttaqien, Tomohiro Motoda, Ryo Hanai et al.· 1 citation
Model Predictive Control (MPC) provides a structured and constraint-aware mechanism for decision-making, but its reliance on optimization-friendly analytical dynamics models limits its use in tasks with contacts and other hard-to-model state dependencies. Model-free reinforcement learning avoids explicit modeling assum...
Recent progress in reinforcement learning with verifiable rewards (RLVR) has highlighted the effectiveness of simple critic-free policy-gradient methods such as Group Relative Policy Optimization (GRPO). In contrast, actor-critic methods rely on learned value functions whose approximation error can introduce bias throu...
Hsiao-Ru Pan, Florent Draye, Bernhard Scholkopf· 0 citations
A framework that uses sampling-based model predictive control (MPC) as scaffolding for real-world dexterous RL, providing structured prior experience and task-directed guidance during learning without human demonstrations or corrective actions is proposed.
Emek Barış Küçüktabak, Karankumar Patel, Zhao-Dong Yang et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.