Results indicate that combining meta-initialization, nonlinear constraint shaping, and topology-aware action masking improves stationary optimization and disturbance recovery within the controlled simulator.
Abstract
Early-exit deep neural networks (DNNs) can reduce edge-inference latency, but abrupt variations in wireless and computing resources can destabilize split-inference policies. This paper proposes a meta-learning-driven adaptive control framework for joint backbone splitting and early-exit routing in MobileViT. The framework formulates multi-exit splitting as a constrained Markov decision process (CMDP) and introduces splitting-aware multi-dimensional adaptive proximal policy optimization (SMAPPO). SMAPPO combines nonlinear quality-of-service (QoS) penalties with topology-aware action masking, while cross-environment meta-initialization supports edge-local adaptation after resource disturbances. Under the stated simulation assumptions, SMAPPO reached the highest performance-index plateau among six methods in a representative 500-episode stationary trace and achieved the lowest normalized total cost across three latency–energy preference settings. Across ten seeds and nine stationary or disturbed scenarios, online SMAPPO achieved a 77.20% measured accuracy and 22.40 mJ of system energy. With an adaptation horizon of K=14, SMAPPO yielded a post-disturbance mean latency of 37.68 ms, a QoS-violation rate of 2.24%, and an on-time completion rate of 98.69%. These results indicate that combining meta-initialization, nonlinear constraint shaping, and topology-aware action masking improves stationary optimization and disturbance recovery within the controlled simulator.
Simulation results confirm that integrating EE with edge computing significantly improves the trade-off between inference accuracy and latency, achieving up to 212% improvement in the average task completion ratio compared to edge computing systems without EE, under the considered simulation settings.
Simone Angelucci, R. Valentini, M. Levorato et al.· Comput. Networks· 0 citations
An online Deep Reinforcement Learning (DRL) based adaptive partition method to dynamically determine optimal partitioning decision so as to jointly accelerate DNN inference and mitigate energy consumption is developed.
Shu-Bin Zhang, Junrong Ma, Kai-Kai Chi et al.· ACM transactions on sensor n...· 0 citations
GMM-TDQN is proposed, a two-stage multi-objective reinforcement learning framework for large-scale edge server deployment that adopts a Transformer-enhanced Deep Q-Network to learn adaptive deployment policies that balance multiple objectives.
Zhou Zhou, Ting-Yu Zheng, Yi-Fu Zeng· Proceedings of the Thirty-Fi...· 0 citations
Quality-of-service (QoS)-aware random access requires adaptive allocation of a finite random access channel (RACH) preamble budget across heterogeneous traffic and access procedures. This paper proposes QP-BD3QN-RACH, a quota-projected branching deep reinforcement learning controller for mixed two-step (2RA) and four-s...
Jiu-ling Guo, Jia-Han Xu, Jia-Shuo Zhang et al.· 0 citations
Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation ac...
Jia-Hua Yang, Zhiwei Yang, Xian-Peng Zhang et al.· 0 citations
The sixth-generation (6G) mobile communication system is anticipated to provide unprecedented performance, but faces a sustainability barrier from overactivated access points (APs). This study addresses energy-efficient operation in the 6G artificial intelligence–radio access network (AI-RAN) by jointly optimizing AP s...
Seonghoon Kim, Dong-eui Kim, Jih-Wan P. Choi· IEEE Transactions on Wireles...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.