Priority-Aware Spectrum Allocation for IoMT-URLLC: A Review of DRL Modeling, Metrics, and PPO/SAC Selection
Abstract
Internet of Medical Things (IoMT) systems supporting remote surgery, ward monitoring, emergency alert systems, and wearable sensors increasingly rely on ultra-reliable low-latency communication (URLLC) transmission. This targeted review systematically summarizes existing research on deep reinforcement learning (DRL)-based priority spectrum scheduling under IoMT-URLLC scenarios. It categorizes relevant standards, medical communication literature, network orchestration papers, and DRL studies from five perspectives: service type, Markov decision process (MDP) modeling, performance metrics, algorithm principles, and practical deployment limits. Proximal policy optimization (PPO) and soft actor-critic (SAC) are further analyzed and compared. When optimization priorities lie in tail-latency guarantee, link reliability, and prevention of critical medical service violations, PPO provides more stable performance. For dynamic monitoring and diagnostic tasks focusing on age of information (AoI), power saving, data reuse, and continuous control, SAC is more suitable, provided that its random exploration range is restricted by predefined medical priority constraints. Current research still has clear limitations: inconsistent experimental benchmarks, insufficient direct comparative studies between PPO and SAC for medical URLLC, and inadequate verification of channel preemption strategies for high-safety medical services.