Skip to content
Preprint

Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry

Aug 2026 · 0 citations · 16 references
Computer Science

Abstract

The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumptions. However, real-world applications often involve heavy-tailed reward distributions and decentralized, information-asymmetric interactions. We study multi-agent multi-armed bandits with heavy-tailed rewards under three information-asymmetry regimes: unobserved actions with common rewards, observed actions with independent rewards, and unobserved actions with independent rewards. We develop robust decentralized algorithms for each setting and derive regret guarantees that nearly match centralized heavy-tailed rates. Experiments on a Pareto-distributed reward environment validate our theoretical findings and illustrate the trade-offs between synchronization, coordination, and exploration across the three regimes.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.