Feedback-Driven Online Deep Reinforcement Learning for Hotspot Aware Adaptive Traffic Steering in Non-Stationary Networks
Modern communication networks increasingly operate under non-stationary traffic conditions, where busty traffic and flash crowds challenge traditional static rule-based network control mechanisms. Despite the fact that reinforcement learning has already been explored for network optimization, most existing methods rely on offline-trained policies that lack stable adaptation to traffic in the network that causes high dimensional state. This paper proposes a feedback-driven online deep reinforcement learning framework for versatile traffic steering in mesh networks. The traffic steering problem is considered as a closed loop evaluation process in which a Deep Q-Network (DQN) constantly updates its policies during runtime. To balance the performance in the network, the framework implements a hotspot-aware lightweight state representation, composed of queue length and link utilization for the top three most congested links alongside end-to-end delay, packet loss rate, and throughput. The proposed framework achieves lower delay and packet loss, faster adaptation, and stable throughput compared to the existing static routing and offline-trained RL policies, while maintaining low monitoring overhead.