Benchmarking QoS-Aware PPO Against Classical Schedulers for 6G Radio Resource Management
Abstract
Future sixth-generation (6G) radio access networks require schedulers that balance throughput, spectral efficiency, fairness, delay, packet loss, and quality-of-service (QoS) satisfaction. This study presents an audited and reproducible benchmark of a QoS-aware Proximal Policy Optimization (QoS-PPO) scheduler against Proportional Fair, Modified Largest Weighted Delay First (M-LWDF), Round Robin, Random Allocation, and Max-CQI. The revised methodology specifies a 60-dimensional state, a factorized five-PRB action, explicit physical-layer and mobility parameters, exact reward normalization, a structured six-candidate hyperparameter search, five independent training seeds, 150,000 training timesteps per seed, matched non-parametric tests, and measured inference latency. The five-seed learning curves plateau within the training budget, but convergence does not yield competitive scheduling performance. Across 48,000 paired evaluation records, Round Robin achieved the highest mean composite reward (0.3647), followed by M-LWDF (0.3425), Max-CQI (0.3371), Random (0.3083), Proportional Fair (0.3023), and QoS-PPO (0.2227). All paired QoS-PPO comparisons remained significant after Holm correction. Median CPU decision time was 440.26 microseconds for QoS-PPO, compared with 3.18 microseconds for Proportional Fair and 7.06 microseconds for M-LWDF. The results provide a conservative benchmark showing that policy convergence alone is insufficient evidence of superiority and that classical schedulers remain strong references for QoS-oriented 6G RRM.