A 6-DOF Deep Reinforcement Learning Architecture for Autonomous Debris Capture and Orbit Transfer
Abstract
While Deep Reinforcement Learning (DRL) offers robust real-time control for Active Debris Removal (ADR), current literature often assumes idealized dynamics and isolated mission segments. This paper presents an end-to-end 6 Degrees of Freedom (6-DOF) Guidance, Navigation, and Control (GNC) architecture using Proximal Policy Optimization (PPO) for a complete four-phase ADR mission. Bypassing standard linear approximations, the custom environment utilizes an absolute Keplerian N-Body Earth-Centered Inertial (ECI) propagator. It enforces operational realism through continuous Tsiolkovsky propellant dynamics, unmitigated kinematic handovers between phases, and instantaneous Huygens-Steiner inertia updates upon target capture. Trained via Curriculum Learning and astrodynamic energy rewards, the agent successfully executes fuel-optimized orbital transfers, performs dynamic obstacle avoidance, stabilizes proximity operations despite handover drift, and autonomously adapts to the severely degraded mass dynamics of the composite orbital vehicle.