Skip to content
Conference

Variance-Reduced Q-Learning over Static and Time-Varying Networks

May 2026 · American Control Conference · pp. 2693-2700 · 0 citations · 21 references
Computer Science Engineering

TL;DR

It is proved that such speedups in sample-complexity require only $\tilde O\left( 1 \right)$ communication, substantially improving upon the communication costs in prior work.

Abstract

We investigate a decentralized reinforcement learning problem involving multiple agents that interact with the same Markov Decision Process (MDP). The agents can exchange information over a network to collectively learn the optimal state-action value function. For this setting, we introduce a novel epoch-based distributed Q-learning algorithm called VRDQ, where within each epoch, agents locally estimate the Bellman optimality operator and diffuse information using a consensus-based protocol. For both static and time-varying networks, we establish high-probability finite-time convergence rates for VRDQ that enjoy linear speedups from collaboration. Crucially, we prove that such speedups in sample-complexity require only $\tilde O\left( 1 \right)$ communication, substantially improving upon the communication costs in prior work.

View source

Similar papers

Preprint Aug 2026

Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning

This work establishes finite-time rates of $\tilde{O} (1/\sqrt{n})$ for the aforementioned two algorithms under asynchronous Markovian sampling, where $n$ is the iteration index and $\tilde{O}$ hides logarithmic expressions.

Ankur Naskar, A. VivekT, Aditya Kumar et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Decentralized Optimal Equilibrium Learning Over Dynamic Networks

This paper studies decentralized learning of socially optimal equilibria in finite normal-form games over dynamic communication networks. Each agent observes only its own realized payoffs, does not know the game a priori, and can communicate only with time-varying neighbors using low-bandwidth messages. We propose netw...

Seref Taha Kiremitci, Muhammed O. Sayin · 0 citations
Preprint Aug 2026

Projection-Free Bandit Online Optimization for Multi-Agent Systems with Dynamic Regret

This paper proposes a distributed bandit online feedback optimization algorithm that relies solely on real-time input-output data and establishes a sublinear dynamic regret bound that depends on a temporal variation measure of system non-stationarity.

Xia Jiang, Lu Liu, Gang Feng · 0 citations
Preprint Sep 2026

Multi-Agent Event-Triggered LQG Control under Shared Communication Constraints

This letter studies event-triggered linear-quadratic-Gaussian (LQG) control for multi-agent systems sharing a communication network with limited per-step capacity. Although the agent dynamics are decoupled, the communication decisions are coupled through the shared network constraint, leading to a constrained multi-age...

Z. Hashemi, Dipankar Maity · 0 citations
Aug 2026

Distributed Online Push-Sum Dual Averaging for Composite Optimization With Communication Delays.

In large-scale network systems, there is a high demand for online optimization, for instance to track time-varying targets in sensor network systems. This article investigates distributed online composite optimization over time-varying, unbalanced multiagent systems with communication delays. Each network node minimize...

Cong Wang, De-Ming Yuan, Qian Ma et al. · 0 citations
#machine learning Preprint Sep 2026

Adapting to Decision-Relevant Non-Stationarity in Decentralized Heterogeneous Bandits

This work introduces Decision-Relevant Fresh Comparison (DRFC), which uses new, balanced samples from all agents to compare arms at the network level and switches only when fresh global evidence indicates that the common best arm has changed, and proves a high-probability dynamic regret bound with no adaptation term de...

Zhao-Jun Peng · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.