Jul 2026· The Journal of Financial Data Science· Vol 8, pp. 203 - 216· 0 citations
TL;DR
The results suggest that adaptive interaction among AI trading systems may create new forms of algorithmic vulnerability relevant for exchanges, trading venues, execution desks, and market-surveillance teams.
Abstract
Modern electronic markets are increasingly populated by adaptive artificial intelligence (AI)–driven trading systems that continuously learn from market outcomes and from one another. This article studies a specific form of algorithmic interaction risk. We simulate multiple market-making agents using tabular Q-learning to optimize quote aggressiveness in the presence of noise traders. Although the agents do not communicate or optimize a joint objective, repeated interaction leads to persistent spread widening, increased dealer profitability, and deterioration in transaction-cost proxies faced by liquidity demanders. Using memory-1 Q-learning agents that compete only on spread, we find that two-dealer markets converge to outcomes that close 78% of the gap between the static Bertrand–Nash equilibrium and joint monopoly, with 97.5% of the joint-action mass on the diagonal (both liquidity providers quoting the same spread). This article develops several observable diagnostics for detecting such behavior, including spread persistence, quote clustering, and forced-deviation response tests. Unlike prior work focused on informed trading and price discovery in Kyle-style environments, our framework emphasizes quote competition, liquidity provision, and execution quality in electronic markets. The results suggest that adaptive interaction among AI trading systems may create new forms of algorithmic vulnerability relevant for exchanges, trading venues, execution desks, and market-surveillance teams.
Automated market makers (AMMs) are a cornerstone of decentralised finance (DeFi). Constant product markets with concentrated liquidity, such as UniswapV3, are now a well-established design. In these markets, liquidity providers (LPs) face a sequential decision problem: they must decide when to rebalance their positions and which price ranges to allocate capital to as market conditions evolve. We formulate dynamic liquidity provision as a stochastic impulse control problem and use reinforcement learning (RL) to solve it, focusing on providing interpretable solutions. We show that learned policies exhibit rich state-dependent behaviour, allocating liquidity according to mispricing, rebalancing costs, uncertainty, inventory exposure, and heterogeneous risk preferences. These behaviours help compress the left tail of the Profit and Loss (PnL) distribution and avoid catastrophic outcomes under high uncertainty. Finally, we benchmark the RL agents against baseline and sophisticated agents from the AMM microstructure literature and analyse their performance.
Georgios Chionas, Charalampos Kleitsikas, Stefanos Leonardos et al.· 0 citations
Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchants and consumers. Yet whether the LLMs behind these agents can price competently in real markets, where customer preferences are hidden, competitors adapt in real time, and demand can shift without warning, has not been systematically tested. We introduce Bazaar, a dynamic sealed-bid benchmark for multi-attribute auction under these conditions. Despite its dynamics, the benchmark is grounded in closed-form customer utilities, enabling exact evaluation. Across 11 frontier LLMs from four providers, the leading agents on customer acquisition (e.g. Gemini 3.1 Pro) are often not the leading agents on profit (e.g. Opus 4.6). The ranking shifts again under demand shocks: agents that learned fastest pre-shock are typically the slowest to revise their beliefs afterwards, while Gemini 3.1 Pro recovers fastest despite not leading on profit. However, even the strongest agent captures less than a third of hindsight-optimal profit, suggesting current LLMs are progressing in agentic commerce but leave substantial headroom.
Shimaa Ahmed, Yiwei Cai, Mohsen Minaei et al.· 0 citations
This paper examines the transformation of financial markets as trading shifts from human-driven to algorithmically dominated activity, with particular focus on large language models (LLMs) as a distinct and increasingly influential category of market participant. We argue that LLMs represent a qualitative shift from traditional algorithmic investing such as high-frequency trading platforms to novel forms of information processing that simultaneously reduce certain information asymmetries while creating new systemic risks. Building on the Grossman– Stiglitz paradox of informationally efficient markets, we demonstrate how the decreasing cost of information processing paradoxically increases market inefficiencies through data mining and the proliferation of false positives. The paper develops new frameworks for understanding markets where price discovery occurs through the interaction of diverse AI architectures, including high-frequency time-series models and natural language processing systems. It then examines the emergence of company-specific semantic factors as sources of alpha. We conclude that although prices reflect the dominant algorithmic interpretation of information, ultra-fast information processing by AI may render markets less predictable rather than more efficient, as prediction itself becomes endogenous to the system being predicted. The paper’s key arguments and findings are as follows:
Agent-based models of markets readily produce emergent instabilities, but telling a genuine collective effect apart from a parameter artefact takes discipline. We apply Bouchaud's phase-diagram method to a continuous-double-auction order-book model. The method is to map the full phase diagram, test its robustness to rule changes, and rule out degenerate and numerical origins before we call any feature a tipping point. The model has fundamental-anchored zero-intelligence liquidity and a mid-anchored chartist herding layer, controlled by the fraction $\varphi$ and the strength $\kappa$ of herders. A 7x6 grid (336 runs, each with a scrambled-sign null) locates an emergent liquidity-stress crossover. The order parameter, the fraction of events with a one-sided book, rises to about 0.34 at $(\varphi,\kappa)=(0.9,1.0)$, is zero across all 42 scrambled cells, and forms a smooth crossover rather than a discontinuous Dark Corner. The dry-up is rule-robust (it recurs under an order-flow-imbalance rule), horizon-robust (about 0.32-0.35 across a 16x range of momentum window), and has a monotone onset boundary $\varphi^*(\kappa) = \{0.55, 0.45, 0.36\}$. We then decompose the mechanism at a matched directional-bias amplitude (mean |p_buy - 0.5| about 0.269). Price-momentum herding carries a large, comparator-robust reflexive component (+0.29; buying begets buying), whereas the order-flow rule's component is about 0 and comparator-dependent. The RMS-mispricing gradient is a placement artefact, largest at $\kappa=0$. A companion two-market analysis finds no directional cross-market contagion across a signal-only herding link.
This work introduces NextFund, an evaluation platform that makes financial-agent behavior observable under live market conditions, and presents NextFund on Hong Kong, U.S., and China A-share equities, illustrating how inspectable decision histories enable fairer benchmarking and more actionable diagnosis.
Changlun Li, Peixian Ma, Qiqi Duan et al.· 0 citations
In over-the-counter corporate bond markets, dealers compete for client trades by quoting bid and ask prices. Tighter quotes attract more business, but also informed customers more likely to trade ahead of adverse price moves, leaving the dealer holding the risk. As dealers increasingly use machine learning to set quotes, they retrain these models on the trades their own quotes attract, creating a feedback loop in which each model reshapes the market that generates its next training data. The question is therefore not only whether a quoting model performs well, but whether the market it creates stays stable as the model learns from it. Existing performative prediction theory gives a sharp stability condition, yet expresses it through abstract properties of the learning objective a trading desk cannot measure before deployment. We introduce REFLEX, a framework that replaces those unobservable quantities with three measurable features of dealer behavior: how strongly trading volume responds to tighter quotes, how sharply the dealer's objective bends around its optimum, and how quickly informed flow increases as spreads narrow. REFLEX combines these into a single retraining modulus, a pre-deployment stability margin estimated from a desk's own quote and execution history that predicts whether repeated retraining will converge or amplify itself. In simulation, predicted and measured stability agree within 8%, and competing dealers increase instability by 1.74x with two and 3.16x with three, as predicted. Where ordinary retraining becomes unstable at modulus 1.21, a structurally anchored correction converges as blind retraining collapses. Calibrated over 36 years of public market data, stability headroom falls roughly 4.4x for investment grade and 4.3x for high yield from calm to crisis regimes. Ultimately, REFLEX turns an abstract convergence theorem into a market-level safety margin.