May 2026· American Control Conference· pp. 2693-2700· 0 citations· 21 references
Computer ScienceEngineering
TL;DR
It is proved that such speedups in sample-complexity require only $\tilde O\left( 1 \right)$ communication, substantially improving upon the communication costs in prior work.
Abstract
We investigate a decentralized reinforcement learning problem involving multiple agents that interact with the same Markov Decision Process (MDP). The agents can exchange information over a network to collectively learn the optimal state-action value function. For this setting, we introduce a novel epoch-based distributed Q-learning algorithm called VRDQ, where within each epoch, agents locally estimate the Bellman optimality operator and diffuse information using a consensus-based protocol. For both static and time-varying networks, we establish high-probability finite-time convergence rates for VRDQ that enjoy linear speedups from collaboration. Crucially, we prove that such speedups in sample-complexity require only $\tilde O\left( 1 \right)$ communication, substantially improving upon the communication costs in prior work.
This work establishes finite-time rates of $\tilde{O} (1/\sqrt{n})$ for the aforementioned two algorithms under asynchronous Markovian sampling, where $n$ is the iteration index and $\tilde{O}$ hides logarithmic expressions.
Ankur Naskar, A. VivekT, Aditya Kumar et al.· 0 citations
This paper studies decentralized learning of socially optimal equilibria in finite normal-form games over dynamic communication networks. Each agent observes only its own realized payoffs, does not know the game a priori, and can communicate only with time-varying neighbors using low-bandwidth messages. We propose netw...
Seref Taha Kiremitci, Muhammed O. Sayin· 0 citations
This paper proposes a distributed bandit online feedback optimization algorithm that relies solely on real-time input-output data and establishes a sublinear dynamic regret bound that depends on a temporal variation measure of system non-stationarity.
This letter studies event-triggered linear-quadratic-Gaussian (LQG) control for multi-agent systems sharing a communication network with limited per-step capacity. Although the agent dynamics are decoupled, the communication decisions are coupled through the shared network constraint, leading to a constrained multi-age...
In large-scale network systems, there is a high demand for online optimization, for instance to track time-varying targets in sensor network systems. This article investigates distributed online composite optimization over time-varying, unbalanced multiagent systems with communication delays. Each network node minimize...
Cong Wang, De-Ming Yuan, Qian Ma et al.· IEEE Transactions on Cyberne...· 0 citations
This work introduces Decision-Relevant Fresh Comparison (DRFC), which uses new, balanced samples from all agents to compare arms at the network level and switches only when fresh global evidence indicates that the common best arm has changed, and proves a high-probability dynamic regret bound with no adaptation term de...
Zhao-Jun Peng· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.