Skip to content

Online Robust Reinforcement Learning Through Monte-Carlo Planning

Sep 2026 · International Conference on Machine Learning · 4 citations · 31 references
Computer Science

TL;DR

A new robust variant of MCTS that mitigates dynamical model ambiguities to bridge the gap between simulation-based planning and real-world deployment and empirical evidence is provided that this method achieves robust performance in planning problems even under significant ambiguity in the underlying reward distribution and transition dynamics.

Abstract

Monte Carlo Tree Search (MCTS) is a powerful framework for solving complex decision-making problems, yet it often relies on the assumption that the simulator and the real-world dynamics are identical. Although this assumption helps achieve the success of MCTS in games like Chess, Go, and Shogi, the real-world scenarios incur ambiguity due to their modeling mismatches in low-fidelity simulators. In this work, we present a new robust variant of MCTS that mitigates dynamical model ambiguities. Our algorithm addresses transition dynamics and reward distribution ambiguities to bridge the gap between simulation-based planning and real-world deployment. We incorporate a robust power mean backup operator and carefully designed exploration bonuses to ensure finite-sample convergence at every node in the search tree. We show that our algorithm achieves a convergence rate of $\mathcal{O}(n^{-1/2})$ for the value estimation at the root node, comparable to that of standard MCTS. Finally, we provide empirical evidence that our method achieves robust performance in planning problems even under significant ambiguity in the underlying reward distribution and transition dynamics.

View source

Similar papers

#machine learning Preprint Sep 2026

Minimax-Optimality of Posterior Sampling for Reinforcement Learning

Posterior sampling for reinforcement learning (PSRL) is one of the simplest and most effective exploration methods, but a basic question has remained open: does unmodified PSRL achieve minimax regret without structural assumptions on the prior? We answer yes. Exact vanilla PSRL is minimax optimal in leading-order Bayes...

T. Goo, Kihyuk Hong · 0 citations
#machine learning Preprint Sep 2026

Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning

This paper proposes a model-based bootstrap framework for uncertainty quantification of OPE in finite-horizon, time-inhomogeneous Markov decision processes (MDPs), and establishes bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation for the target policy va...

Wei-Wei Wang, Yu-Qiang Li, Xian-Yi Wu et al. · 0 citations
#artificial intelligence Conference Sep 2026

Power Mean Estimation in Stochastic Continuous Monte-Carlo Tree Search

A novel MCTS algorithm, \Algname, designed for continuous, stochastic MDPs, that integrates a power mean as a value backup operator, alongside a polynomial exploration bonus to address the non-stationarity inherent in continuous action spaces.

T. Dam · 3 citations
#small language model Preprint Aug 2026

Test-time Reinforcement Learning in Imperfect Information Games

This work extends the concept of gadget game, tabular technique for test-time search, to the reinforcement learning setting and formally proves that, unlike prior tabular algorithms, regularized policy-gradient algorithms limit possible strategy degradation caused by test-time reasoning, even without the gadget games.

Ondrej Kubícek, Viliam Lisý, Tuomas Sandholm · 0 citations
#machine learning Preprint Sep 2026

Near-Optimal Reinforcement Learning with Multi-Step Transition Lookahead

The problem remains hard for every discount factor, and for every fixed discount factor, exact planning remains NP-hard, and a randomized polynomial-time approximation scheme is introduced for every fixed look-ahead depth.

Corentin Pla, Hugo Richard, Marc Abeille et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning

Empirically, EPIG reduces gradient MSE in cloned-state control, winning in all nine dense continuous-control environments of a 13-environment sweep and recovering the reference gradient direction near-perfectly, and it improves frozen-LLM gradient calibration relative to entropy branching.

Nikita Khomich, L. Hermansson, Ido Hakimi · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.