Skip to content
Preprint

Lipschitz Bandits with Arbitrary Feedback Delays

Aug 2026 · 0 citations · 18 references
Computer Science

TL;DR

The authors' bounds match existing delay-free regret guarantees for Lipschitz bandits and characterize the additional $\tilde{O}(\sqrt{D})$ impact introduced by feedback delays.

Abstract

The Lipschitz bandit problem extends the traditional multi-armed bandit framework to continuous action spaces by assuming that the reward functions satisfy a Lipschitz condition. This work investigates Lipschitz bandits under arbitrary feedback delays, where reward signals are not received immediately upon taking an action but after an arbitrarily chosen delay. We consider both stochastic and adversarial reward settings, proposing an elimination-based algorithm and an EXP3-based algorithm, respectively. For both settings, our algorithms achieve a regret bound of $\tilde{O}\left(T^{\frac{d_z+1}{d_z+2}}+\sqrt{D}\right)$ over a time horizon $T$ with total delay $D$, where the main difference between settings lies in the definition of the zooming dimension $d_z$. Our bounds match existing delay-free regret guarantees for Lipschitz bandits and characterize the additional $\tilde{O}(\sqrt{D})$ impact introduced by feedback delays.

View source

Similar papers

Preprint Aug 2026

Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits

Motivated by decentralized applications, we study cooperative multi-agent bandits in continuous (Lipschitz) action spaces when the Lipschitz constant is unknown. We consider three information structures: (A)~unobserved actions with common rewards, (B)~observed actions with independent rewards, and (C)~unobserved action...

Ricardo Parada, Chenzhang Zhao, William Chang · 0 citations
#machine learning Preprint Aug 2026

Constant Individual Regret in General Games

This work introduces \emph{ECHO-OFTRL}: optimistic follow-the-regularized-leader (OFTRL) equipped with an EMA cascade for high-order optimism (ECHO), where EMA denotes exponential moving average, and leverages a new form of optimism inspired by modern filter design.

Mingyang Liu, Gabriele Farina, A. Ozdaglar · 5 citations · ⚡4
#machine learning Preprint Oct 2026

Rate-Optimal Algorithm for Adversarial Linear CMDPs

We study episodic adversarial linear constrained Markov decision processes (CMDPs) with unknown transitions, where both the loss and constraint functions may vary adversarially across episodes. The best previous algorithm achieves $\widetilde{\mathcal{O}}(K^{3/4})$ regret and cumulative constraint violation, leaving a...

Kihyun Yu, Hong-Hao Wei, Dabeen Lee · 0 citations
#machine learning Preprint Sep 2026

Learning Infinite-Horizon Average-Reward CMDPs via State Augmentation

We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Existing high-probability guarantees for this setting either require computationally inefficient algorithms or have suboptimal dependence on the number of interactions $T$. We propose, to th...

Kihyun Yu, Seoungbin Bae, Dabeen Lee · 0 citations
#machine learning Preprint Aug 2026

Memory--Batch Tradeoffs in Lipschitz Bandits

The analysis separates fine-scale comparisons from the regional information needed to allocate their samples at low regret, which separates fine-scale comparisons from the regional information needed to allocate their samples at low regret.

Zicheng Lyu, Zengfeng Huang · 0 citations
#artificial intelligence Preprint Sep 2026

Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivity

We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multiple correct answers. For $K$-armed bandits with $A$ optimal arms, we first provide a sharper analysis of previous sub-sampling algorithms (De Heide et al., 2021; Zhu and Nowak, 20...

Kaixuan Ji, Qi-Wei Di, Qing-Yue Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.