Skip to content

Regime Structure in Adaptive Tariff Conflicts: A Multi-Agent Reinforcement Learning Analysis

Sep 2026 · IEEE Conference on Computational Intelligence for Financial Engineering & Economics · pp. 122-131 · 0 citations · 14 references

Abstract

Adaptive strategic conflicts unfold through repeated, history-dependent interactions that conventional equilibrium analysis captures only imperfectly. We present a multi-agent reinforcement learning (MARL) phase-diagram methodology that maps structural parameter spaces to stable outcome regimes and coordination-sensitive regions, demonstrated on a stylized two-country tariff conflict in which a structurally privileged country deploys tariffs and currency depreciation and a responding country sets counter-tariffs and competitive depreciation, with both agents updating simultaneously rather than under explicit Stackelberg commitment. Agents learn via tabular Q-learning, with the follower using double-Q updates to reduce overestimation bias on its larger joint action space; this yields co-adaptive policies without imposing equilibrium assumptions and keeps learned value functions directly inspectable. Elasticities and policy ranges are calibrated to the 2018–2019 US–China episode (import demand elasticities of 1.3–1.55, tariff ceilings of 25%–35%), with results reported across an 8×8 grid and five seeds per cell. We identify three regimes (Deterrence, Transition, Escalation) and three principal findings. First, the Deterrence–Escalation boundary tracks the ratio of the follower’s retaliation ceiling to the leader’s tariff ceiling. Second, leader currency depreciation extends Deterrence only at moderate intensity, with a non-monotone inverted-W effect including an interior optimum and partial recovery at higher depreciation levels. Third, diplomatic cost and export sensitivity jointly compress the Deterrence region. Near regime boundaries, identical parameters reach different stable outcomes across seeds, revealing coordination-sensitive regions with multiple attractors rather than learning noise. The methodology generalizes to repeated strategic interactions with parameter-sensitive outcomes.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.