Skip to content
Open access

Exploration Is a State, Not a Setting: A Markov-Switching Reinforcement-Learning Model of Strategy Transitions in Sequential Choice

Aug 2026 · Journal of Computing Theories and Applications · 0 citations · 33 references

TL;DR

It is concluded that exploration is better described as a dynamic state than as a fixed trait, and that modeling it as a switching process is both more accurate and useful.

Abstract

Computational accounts of human exploration usually assume a stationary policy, in which a single set of parameters for value sensitivity and uncertainty seeking generates every choice in a task, so that all within-task variation is treated as decision noise. Neuroscience instead treats exploration and exploitation as dissociable modes between which the brain switches, which a stationary model cannot represent. We introduce the Markov-Switching Reinforcement-Learning (MS-RL) model, in which a first-order hidden Markov chain governs transitions among a small number of latent decision regimes, each with its own softmax policy over a shared value-learning process. The three regimes are exploitation, directed exploration, and random exploration. The model contains the standard stationary account (one regime) and a temporally unstructured mixture (memoryless transitions) as nested special cases, making the stationarity assumption testable. We estimate the model using Expectation-Maximization with Viterbi decoding and select the number of regimes using the Bayesian information criterion. A parameter- and state-recovery study confirmed that the generating parameters and latent regime paths are recoverable (parameter correlations 0.88–0.94; state accuracy 86 percent; Cohen’s kappa ≈ 0.79). Applied to an openly available two-armed bandit dataset (46 adults, 13,800 choices), a three-regime MS-RL model was preferred over the nested baselines, two- and four-regime variants, and a single-regime model with smoothly time-varying value sensitivity. An ablation analysis attributes the largest gains to the latent regimes and temporal switching. The decoded regime path revealed a systematic within-game shift from directed exploration toward exploitation. The estimated transition matrix provides a per-participant measure of strategy change that has no counterpart in stationary models. We conclude that exploration is better described as a dynamic state than as a fixed trait, and that modeling it as a switching process is both more accurate and useful.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning

Biological agents do not learn under conditions of unlimited computation. For humans, learning and choice are shaped by constraints on perception, attention, and working memory, which limit how much state information guides behavior and therefore bound policy complexity. Standard reinforcement learning models typically...

James Wu, C. Sims · 0 citations
Preprint Sep 2026

Transformation of Adaptive Multistage Sampling for Solving Finite-Horizon Markov Decision Processes with Unknown Model

It is shown that AMR is asymptotically optimal such that the sequence of the expected absolute errors approaches zero and its convergence rate depends on the number of visits to each reachable state at each stage from the initial state, essentially transforming the result of AMS into the RL setting.

H. Chang · 0 citations
#machine learning Preprint Sep 2026

Temporal-Causal Inference for Reinforcement Learning via Automata Learning

We consider reinforcement learning in environments with dynamics that undergo an irreversible phase transition governed by a hidden temporal pattern. The agent observes the base state but cannot observe the phase directly. We formalize this problem as a two-phase non-Markovian decision process and introduce Temporal-Ca...

Jan Corazza, Daniil Kaminskyi, Simon Lutz et al. · 0 citations
Open access Aug 2026

Embodied Learning under Policy and Dynamics Shifts

This work proposes Transition Occupancy Matching as a unifying principle to resolve policy and dynamics shifts within a single mathematical framework and introduces Occupancy-Matching Policy Optimization (OMPO), a novel algorithm that optimizes a surrogate objective explicitly correcting for transition discrepancies.

Yu Luo, Lei Lv, Fu-Chun Sun et al. · 0 citations
Preprint Aug 2026

Learning to Control Coupled-Dynamics Environments with Joint Markov Decision Processes

A nonparametric distributional Bellman optimality operator for JMDPs is defined, and it is proved that when the induced marginal MDP has a unique optimal policy, its iterates converge in Wasserstein distance to the optimal joint return law.

Ege C. Kaya, Aliasghar Pourghani, Mahsa Ghasemi et al. · 0 citations
#machine learning Preprint Sep 2026

Self-Confirming Superposition Traps in Reinforcement Learning

It is shown that this loop can sustain a lower-return policy even when representation fitting is globally optimal on data selected by the agent, which then uses the resulting returns to guide its next choices.

Dai Shi, Andi Han, Feng Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.