This paper proposes a distributed bandit online feedback optimization algorithm that relies solely on real-time input-output data and establishes a sublinear dynamic regret bound that depends on a temporal variation measure of system non-stationarity.
Abstract
This paper investigates distributed online optimization for multi-agent dynamical systems with constrained inputs and time-varying cost functions. While online convex optimization offers a principal framework for sequential decision-making, existing online learning and optimization algorithms typically require accurate system models, limiting their applicability in practical settings. To overcome this challenge, we propose a distributed bandit online feedback optimization algorithm that relies solely on real-time input-output data. The algorithm employs a smoothing zeroth-order one-point estimator to construct local gradient approximations directly from cost evaluations. Additionally, to enforce input constraints effectively, we integrate a projection-free conditional gradient update, making the algorithm well-suited for online and large-scale settings. Furthermore, we establish a sublinear dynamic regret bound that depends on a temporal variation measure of system non-stationarity. Finally, numerical simulations demonstrate the effectiveness of the proposed algorithm.
A hybrid offline-online multi-agent reinforcement learning framework based on decision transformers that incorporates return-weighted sampling, a critic conditioned on neighbors' actions, and neighborhood-correlated exploration that achieves quality-of-service (QoS) performance comparable to centralized methods.
Yi-Ming Zhang, Kun Yang, Cong Shen et al.· 0 citations
This article analyzes the optimal containment control problem of discrete-time multiagent systems (MASs). Multistep temporal difference (TD) learning is integrated with policy gradient (PG) reinforcement learning (RL) to form an online off-policy multistep PG (MS-PG) algorithm. The proposed MS-PG algorithm achieves opt...
Kai-Tian Chen, Huai-Cheng Yan, Qi-Wei Liu et al.· IEEE Transactions on Cyberne...· 0 citations
A randomized two-point direct-search algorithm for nonconvex time-varying optimization and derive iteration-complexity bounds under both constant and diminishing probing ratios, which recover the complexity of existing zeroth-order methods in the time-invariant setting while extending direct- search methods beyond stat...
Model misspecification poses a fundamental challenge in multi-agent reinforcement learning, where transition uncertainty can be amplified by strategic interactions among agents. Distributionally robust Markov games (DRMGs) provide a principled framework for addressing such uncertainty, yet existing methods often rely o...
Debamita Ghosh, George Atia, Yue Wang· 0 citations
This paper studies real time formation control for large scale multi agent systems (LMAS) with anonymous agents and finite time requirements. Instead of solving a centralized Hamilton Jacobi Bellman (HJB) problem or a coupled mean field game system, we design an assignment free controller directly at the distribution l...
Nishad Tasnim, Ze-Jian Zhou· Conference on Control Techno...· 0 citations
A line-search-free and function-value-free adaptive projected-gradient algorithm for the sample-average approximation (SAA) problem that transfers vanishing SAA residuals to Pareto stationarity for the population problem, while an additional concentration argument gives a finite-sample residual bound on compact sets.
Yi-Yang Li, Lei Wang, Xiaojun Chen· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.