Skip to content
Preprint

Data-Driven Brownian Reflection Control

Sep 2026 · 0 citations
Mathematics

Abstract

We study a data-driven reflection control problem for a Brownian model with unknown drift and volatility. We first propose a learn-then-optimize (LTO) algorithm: it estimates the policy-relevant parameter during exploration, plugs the estimate into the optimality equation, and exploits the resulting policy---achieving an $O(\sqrt{T})$ finite-time expected regret bound. We further propose two algorithms, adaptive-updating (AU) and full-history adaptive-updating (AU-FH), which continuously update the estimator and reflecting level, attaining an improved $O(\log T)$ regret bound. Notably, AU-FH algorithms leverages all historical data, yielding better performance in numerical simulations. Our analysis decomposes regret into exploration, transient, and learning components. Transient regret from nonstationarity is bounded by the time-integrated deviation of the transition semigroup from stationarity evaluated on the holding cost, which can be further bounded via a Foster-Lyapunov inequality for exponential convergence of the controlled reflected Brownian motion (RBM). For learning regret, we establish local regularity properties together with consistency and mean-squared error bounds for the estimator, which control the stationary cost gap between the learned and optimal reflection policies. In addition, we leverage the monotonicity of the moving boundary Skorokhod map to derive moment bounds for the AU algorithms'switching states, via pathwise comparison with fixed-boundary RBMs.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.