Skip to content
Conference Open access

Constant-Memory Strategies in Stochastic Games: A Theoretical and Empirical Study

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · 0 citations · 56 references

TL;DR

This work comprehensively investigates the concept of constant-memory strategies in stochastic games, and uncovers the connection between decision models in single-agent and multi-agent contexts.

Abstract

Stochastic games have become a prevalent framework for studying long-term multi-agent interactions, especially in the context of multi-agent reinforcement learning. In this work, we comprehensively investigate the concept of constant-memory strategies in stochastic games. We first establish some results on best responses and Nash equilibria for behavioral constant-memory strategies, followed by a discussion on the computational hardness of best responding to mixed constant-memory strategies. Those theoretic insights are later verified on several sequential decision-making testbeds, including the Iterated Prisoner's Dilemma, the Iterated Traveler's Dilemma, and the Pursuit domain. This work aims to enhance the understanding of theoretical issues in single-agent planning under multi-agent systems, and uncover the connection between decision models in single-agent and multi-agent contexts. The codebase and the full version of this paper is available at github.com/Fernadoo/Const-Mem.

Read PDF

Similar papers

Preprint Aug 2026

Planning Against Learning in Rank-1 Games

Learning algorithms are often used to make decisions in repeated multi-agent environments. When another player understands how a learner adapts from past experience, that player can plan strategically across rounds to influence the learner's future behavior. Recent work shows that optimizing against Replicator Dynamics...

William Overman · 0 citations
Open access Sep 2026

Research on the Game Dynamics of Optional Public Goods

The main findings demonstrate that individuals update strategies primarily through self-adjustment based on historical payoffs, with imitation playing merely an auxiliary role and the optimal self-adjustment proportion is approximately 0.9, and low sensitivity coefficients and mutation rates favor the emergence of coop...

Hao-Chen Wu, Meng-Cheng Sun, Lu-He Yang et al. · 0 citations
Preprint Aug 2026

Equilibrium in Multi-Agent Reinforcement Learning

Standard solution concepts for stochastic games, such as Markov perfect equilibrium and Markov coarse correlated equilibrium, are computationally difficult, and thus, standard decentralized reinforcement-learning algorithms should not generally be expected to converge to them. In this paper, we study the equilibrium ge...

Maurizio D'Andrea, Bar Light · 0 citations
Preprint Sep 2026

Absorbing State Phase Transitions in Multi-Agent Search

Nontrivial dynamics can emerge in large language model (LLM)-based multi-agent systems, and preliminary evidence exists that formalisms from statistical mechanics can be effective at modeling and predicting such behaviors. In parallel, designing multi-agent communication topology for optimal task-solving is an active r...

Wen-Wen Zheng, Yuzhe Yang, Helen Qu et al. · 0 citations
Preprint Aug 2026

What preferences can - and cannot - predict in multi-agent online learning

A three-player game is constructed with a preferentially stable set whose span is dynamically unstable, showing that preferences do not suffice as a criterion of dynamic stability and bridges the gap via the notion of resilience under aggregate deviations.

Omar Abbadi, R. Laraki, P. Mertikopoulos · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.