Skip to content

Similar papers

#artificial intelligence Preprint Aug 2026

Hybrid Offline-Online Multi-Agent Decision Transformers for Wireless Resource Management

This paper develops a hybrid offline-online multi-agent reinforcement learning framework based on decision transformers. The policy is first pretrained offline via supervised sequence modeling of trajectories generated by existing policies, providing a safe and sample-efficient initialization. It is then fine-tuned online using a hybrid objective that incorporates critic-guided gradients, enabling performance improvements beyond the offline policy. To facilitate stable offline-to-online transfer and effective multi-agent coordination, the framework incorporates return-weighted sampling, a critic conditioned on neighbors'actions, and neighborhood-correlated exploration. The approach is fully distributed: both training and execution rely only on local observations and limited information exchange among neighboring agents. Evaluations with dynamic traffic arrivals in two settings: (i) joint scheduling and power allocation and (ii) coordinated beamforming, show that the proposed method achieves quality-of-service (QoS) performance comparable to centralized methods. Moreover, when pretrained on lower-quality datasets, online fine-tuning is also observed to surpass the initial offline policy. These results demonstrate a promising learning-based alternative for wireless resource management.

Yiming Zhang, Kun Yang, Cong Shen et al. · 0 citations
Conference Jul 2026

Joint AoI and SWIPT-Aware Scheduling via Multi- Agent Deep Reinforcement Learning

This work investigates the joint optimization of Age of Information (AoI) and energy harvesting (EH) in wireless edge computing systems, where edge servers not only process IoT data but also act as wireless power suppliers via simultaneous wireless information and power transfer (SWIPT). Building upon the asynchronous model-free fractional multi-agent reinforcement learning framework and the Lyapunov drift-plus-penalty (DPP) concept, we design a fractional-based reward function for AoI and construct a virtual queue to enforce long-term energy stability under battery storage constraints. The overall reward is formulated as a weighted sum, capturing the trade-off between timeliness and energy sustainability, with update decisions, task offloading, and power splitting ratios as key control variables. Simulation results demonstrate that the developed multi-agent deep reinforcement learning approach achieves superior AoI–energy trade-offs compared to related baseline algorithms. These findings highlight the effectiveness of our framework in balancing information freshness and sustainable energy harvesting under resource-constrained edge environments.

Kuang-Ting Liu, Jain-Shing Liu, Wan-Ling Chang · 0 citations
Preprint Jul 2026

Value-Aware Prediction for Robust Multi-Agent Coordination Under Communication Loss

A value-aware extension of Multi-Agent Observation Sharing under Communication Dropout to patch communication gaps is proposed; it is referred to as Value-Aware MARO and dynamically weighting the predictor's loss function using advantage estimates derived from the underlying actor-critic architecture.

K. D. Kafadar, Eren Özaltun, M. E. Şanlı et al. · 0 citations
Jul 2026

Safe Multi-Agent Collaborative Learning for Networked Grid Operation Under Power Network Coupling Constraints

This paper formulate networked grid operation as a constrained decentralized partially observable Markov decision process and proposes a safe multi-agent collaborative learning framework that aims to reduce operating cost, load shedding, renewable curtailment, and carbon-relevant corrective burden.

Jiayi Zhang, Bing Fang, Huanxiu Xiao et al. · 0 citations
Preprint Aug 2026

DER Allocation without Load Prediction via Reinforcement Learning

A forecast-free reinforcement learning (RL) framework for DERA allocation that learns optimal policies directly from operational data, which preserves the interpretability and constraint satisfaction of DER model while adapting to stochastic demand variations through data-driven updates.

Abed AlRahman Al Makdah, Aravind Ramana, Shaofeng Zou et al. · 0 citations