GMM-TDQN is proposed, a two-stage multi-objective reinforcement learning framework for large-scale edge server deployment that adopts a Transformer-enhanced Deep Q-Network to learn adaptive deployment policies that balance multiple objectives.
A two-timescale multi-layer deep reinforcement learning framework with a latent action space (2T-MDRL-LA) to jointly optimize service placement, user association, computational delegation, task offloading, and user transmit power and achieves near-optimal performance compared to branch-and-bound solutions.
V. Son, Van-Dinh Nguyen, Ngoc Hung Nguyen et al.· 0 citations
STQ-Scheduler is proposed, a secure and throughput-aware deep reinforcement learning framework that integrates high-throughput data processing, Transformer-based QoE prediction, and Proximal Policy Optimization-based resource scheduling to ensure data quality and prevent data processing from becoming a bottleneck in distributed training.
Yi-Chun Chang, Min-Wei Jiang· Fundamental Scientific Repor...· 0 citations
The proposed modified Deep Reinforcement Learning-based intelligent TDD configuration framework for adaptive radio resource allocation in 5G HetNets effectively enhances network reliability, resource utilization, and communication efficiency in dynamic 5G HetNet environments.
G. Dalton, ·. A. Bamila, Virgin Louis et al.· Wireless networks· 0 citations
Sensitivity and ablation studies confirm stable learning and controllable latency-cost trade-offs, demonstrating that lightweight RL can effectively deliver cost-efficient, adaptive autoscaling in hybrid cloud environments.
Bekzat Kobei, N. Seilova, Zarina A. Kashaganova· AI@DTESI· 0 citations
This paper proposes an edge-cloud collaborative physics-informed reinforcement learning framework for production data center HVAC control that integrates a physics-informed cold-start solution using Adaptive Particle Swarm Optimization, a three-time-scale edge–cloud architecture, and a constraint-aware safe projection layer that embeds thermal safety hard constraints directly into the neural network policy.
Shichao Huang, Yi-Bing Zhou, Yuan Liu· Italian National Conference...· 0 citations
Experimental evaluation on a heterogeneous synthetic benchmark demonstrates that the proposed DDQN scheduler reduces SLA violations by approximately 85% relative to Round Robin and 72% relative to the greedy baseline, while achieving superior energy efficiency.
Vishakha Makode, Taresh Ayaspure· Journal of Advances in Devel...· 0 citations