Skip to content

SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes

Sep 2026 · 0 citations · 99 references
Computer Science

TL;DR

FARSIGHT (Financial Agent Robustness and Security Investigation and Global Holistic Testing), a framework that performs scheme-level evaluation of financial LLM agents on two axes: robustness under market turbulence and security against three attack types: attacks on information sources, attacks on agents, and agent-as-attacker behaviors.

Abstract

Autonomous large language model (LLM) agents are moving rapidly into high-stakes domains, yet existing agentic-AI security studies remain largely domain-agnostic and overlook the distinctive, high-consequence attack surface such settings create. We examine this gap through financial trading agents, a representative case of high-stakes agentic security, where a single compromised agent has direct execution authority over real capital in an adversarial, reflexive market. To this end, we present FARSIGHT (Financial Agent Robustness and Security Investigation and Global Holistic Testing), a framework that performs scheme-level evaluation of financial LLM agents on two axes: robustness under market turbulence (including flash-crash-like scenarios), and security against three attack types: attacks on information sources, attacks on agents, and agent-as-attacker behaviors. Applying FARSIGHT to 15 representative academic schemes, we find that most overlook robustness and realistic adversarial threats: 80% fail at least one core robustness metric and 100% exhibit security vulnerabilities. These two failure modes are inseparable: a small misjudgment can cascade into a market-wide crash on its own, while an adversary can deliberately trigger the same collapse at minimal cost.

View source

Similar papers

#artificial intelligence Preprint Open access Sep 2026

Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems

Multi-agent trading systems built on large language models (LLMs) are beginning to appear in quantitative finance, yet their robustness to adversarial inputs is largely unknown. We study the vulnerability of LLM trading stacks to black-box, input-only attacks that enter solely via admissible social-media feeds. We intr...

Rong-Sua Qi, Jun-Hao Dong, Thai Duc Nguyen et al. · 0 citations
Preprint Aug 2026

Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems

This paper presents the first systematic empirical study in the financial domain to characterize how an adversarial signal enters a multi-agent trading system and how far it survives toward the decision, and evaluates four communication topologies under data- and agent-level attacks.

CheolWon Na, Hao Ni, Lukasz Szpruch et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Your Agent Says Yes: Interpreting Adversarial Market Behavior Beyond Individual Transactions

Repeated runs show that category-level and within-trajectory relations can recur even when normalized score-change rankings do not, and motivate agent-behavior evaluation that links communication, authorization, and evolving state instead of treating individual transaction verdicts as complete safety judgments.

Ze-Lin Li, Yi-Yun Su, Matt White et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures

Reidual attack success concentrates where flags are unparseable or auditors leak; an RL-trained enforcement controller handles hedged and malformed verdicts that rule-based parsing cannot, cutting ambiguous-critique failure to a fraction of the rule-based baseline.

Yu-Hang Wang · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.