Parallel Training of a Yie Ar Kung-Fu Agent Based on Proximal Policy Optimization
Abstract
Improving sample efficiency and achieving balanced multi-task performance across different tasks remain important research challenges in reinforcement learning. Yie Ar Kung-Fu is formulated as a multi-task reinforcement learning environment. A task-specific neural network architecture, simplified state and action encoding schemes, a reward mechanism, and a dynamic learning-rate schedule are designed for the game environment. Based on Proximal Policy Optimization (PPO), a parallel training method is further proposed. The method aggregates episode samples generated from different opponent environments into a unified sample queue in the order of episode completion and dynamically adjusts the weights of samples from different opponents in the loss function based on their win rates. Experiments were conducted on the AutoDL platform using an NVIDIA RTX 4090 GPU. The proposed method enables joint training across multiple opponent environments. After approximately 15,000 training episodes, the trained agent achieved win rates exceeding 90% against all opponents and attained a level-clear rate of 79%, demonstrating strong game-playing performance and balanced performance across multiple opponent tasks. Furthermore, interdisciplinary analogies are used to interpret how parallel training improves sample efficiency and balanced performance across tasks. The proposed method is simple to implement, stable during training, and readily applicable to other complex multi-task environments.