Skip to content
Preprint

Projection-Free Bandit Online Optimization for Multi-Agent Systems with Dynamic Regret

Aug 2026 · 0 citations · 37 references
Computer Science

TL;DR

This paper proposes a distributed bandit online feedback optimization algorithm that relies solely on real-time input-output data and establishes a sublinear dynamic regret bound that depends on a temporal variation measure of system non-stationarity.

Abstract

This paper investigates distributed online optimization for multi-agent dynamical systems with constrained inputs and time-varying cost functions. While online convex optimization offers a principal framework for sequential decision-making, existing online learning and optimization algorithms typically require accurate system models, limiting their applicability in practical settings. To overcome this challenge, we propose a distributed bandit online feedback optimization algorithm that relies solely on real-time input-output data. The algorithm employs a smoothing zeroth-order one-point estimator to construct local gradient approximations directly from cost evaluations. Additionally, to enforce input constraints effectively, we integrate a projection-free conditional gradient update, making the algorithm well-suited for online and large-scale settings. Furthermore, we establish a sublinear dynamic regret bound that depends on a temporal variation measure of system non-stationarity. Finally, numerical simulations demonstrate the effectiveness of the proposed algorithm.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Hybrid Offline-Online Multi-Agent Decision Transformers for Wireless Resource Management

A hybrid offline-online multi-agent reinforcement learning framework based on decision transformers that incorporates return-weighted sampling, a critic conditioned on neighbors' actions, and neighborhood-correlated exploration that achieves quality-of-service (QoS) performance comparable to centralized methods.

Yi-Ming Zhang, Kun Yang, Cong Shen et al. · 0 citations
Aug 2026

Optimal Containment of Multiagent Systems With Multistep Policy Gradient Reinforcement Learning.

This article analyzes the optimal containment control problem of discrete-time multiagent systems (MASs). Multistep temporal difference (TD) learning is integrated with policy gradient (PG) reinforcement learning (RL) to form an online off-policy multistep PG (MS-PG) algorithm. The proposed MS-PG algorithm achieves opt...

Kai-Tian Chen, Huai-Cheng Yan, Qi-Wei Liu et al. · 0 citations
Preprint Aug 2026

Direct Search Methods for Online Nonconvex Optimization Under Inexact Bandit Feedback

A randomized two-point direct-search algorithm for nonconvex time-varying optimization and derive iteration-complexity bounds under both constant and diminishing probing ratios, which recover the complexity of existing zeroth-order methods in the time-invariant setting while extending direct- search methods beyond stat...

Gaspar Robert, Gianluca Bianchin · 0 citations
#artificial intelligence Preprint Sep 2026

Interactive Distributionally Robust Multi-Agent Learning with General Function Approximation

Model misspecification poses a fundamental challenge in multi-agent reinforcement learning, where transition uncertainty can be amplified by strategic interactions among agents. Distributionally robust Markov games (DRMGs) provide a principled framework for addressing such uncertainty, yet existing methods often rely o...

Debamita Ghosh, George Atia, Yue Wang · 0 citations
Conference Aug 2026

Assignment-Free Real-time Tracking Control for Large-Scale Multi-Agent Systems

This paper studies real time formation control for large scale multi agent systems (LMAS) with anonymous agents and finite time requirements. Instead of solving a centralized Hamilton Jacobi Bellman (HJB) problem or a coupled mean field game system, we design an assignment free controller directly at the distribution l...

Nishad Tasnim, Ze-Jian Zhou · 0 citations
Preprint Sep 2026

An Adaptive Projected-Gradient Algorithm for Sample-Average Approximations of Stochastic Multi-Objective Optimization

A line-search-free and function-value-free adaptive projected-gradient algorithm for the sample-average approximation (SAA) problem that transfers vanishing SAA residuals to Pareto stationarity for the population problem, while an additional concentration argument gives a finite-sample residual bound on compact sets.

Yi-Yang Li, Lei Wang, Xiaojun Chen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.