Skip to content

Author

Marco Herrera

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#reinforcement learning Open access Sep 2026

Evaluating Deep Actor–Critic Methods for Path Planning of Mobile Manipulators Under Wheel–Terrain Interaction

Reinforcement learning (RL) has become an effective paradigm for enabling autonomous robots to acquire navigation policies directly from interaction with complex and uncertain environments. Nevertheless, autonomous path planning for Skid-Steer Mobile Manipulators (SSMMs) remains a challenging problem because it requires the coordinated control of the non-holonomic mobile base and the manipulator while simultaneously accounting for obstacle avoidance and wheel–terrain interaction effects. This paper presents and evaluates RL-based path planning strategies for SSMMs, explicitly incorporating coupled dynamics of the mobile platform and manipulator to generate collision-free trajectories under varying terrain conditions. The proposed framework incorporates a slip-aware reward formulation that penalizes discrepancies between commanded and measured robot motion while accounting for longitudinal and lateral slip resulting from wheel–terrain interaction. The main contributions are i) a unified RL-based framework based on actor–critic techniques for SSMM path planning, integrating the mobile base and manipulator dynamics within a coupled system representation; ii) a physics-aware multi-objective reward formulation that incorporates wheel–terrain interaction into policy learning; and iii) the implementation via simulation and field validation of the proposed policies under progressively complex navigation conditions and real underground mining scenarios. The framework is evaluated using four RL algorithms across multiple environments and maps from real mining scenarios, encompassing diverse navigation conditions and start-to-goal configurations. The evaluated methods include Deep Deterministic Policy Gradient (DDPG), Proximal Policy Optimization (PPO), Soft Actor–Critic (SAC), and Twin Delayed DDPG (TD3). Experimental field results show that SAC achieves the lowest planning time, reducing the planning time by 127.3%, 24.1%, and 2.52% compared with PPO, TD3, and DDPG, respectively. SAC also achieves the shortest path, reducing the average path length by 20.32%, 8.58%, and 1.90% compared with PPO, DDPG, and TD3, respectively. Moreover, SAC generates smoother control profiles for both the mobile base and the manipulator arm, while TD3 exhibits competitive performance across several navigation metrics. The proposed framework demonstrates the potential of slip-aware RL for coordinated SSMM navigation, providing a practical foundation for improving the safety, energy efficiency, and operational autonomy of mobile manipulators exposed to complex mining environments.

Christian Camacho, Óscar Camacho, Marco Herrera et al. · 0 citations