Controlled evidence ablations show that recent trading history strongly governs activity prediction, whereas asset selection is substantially more sensitive to the available evidence.
Abstract
Large language models (LLMs) are increasingly used as user simulators, yet it remains unclear whether their predictions faithfully reproduce the evolving decisions of individual users. We investigate this question in a controlled longitudinal paper-trading study with 80 participants, where user interactions, simulated transactions, virtual portfolio states, and point-in-time market information are aligned under a rolling next-day prediction protocol. We evaluate behavioral fidelity hierarchically, from trade occurrence to action structure, asset selection, and downstream portfolio consequences. Across 1,239 aligned user-days, no evaluated LLM reliably outperforms a simple recent-activity persistence baseline for predicting whether a user trades. Fidelity further deteriorates at finer levels: models struggle to recover buy--sell structure and traded assets, and similar activity-level predictions can lead to substantially different portfolio trajectories. Controlled evidence ablations show that recent trading history strongly governs activity prediction, whereas asset selection is substantially more sensitive to the available evidence. An observational analysis further finds that intensified ticker-specific research predicts imminent trading, but diagnostic tests do not support a causal interpretation. These findings suggest that current LLMs capture useful short-term behavioral regularities without yet recovering a stable individual decision mechanism.
Large language models (LLMs) and multi-agent systems (MAS) have shown promise in financial decision-making, yet existing evaluations focus on equity trading and primarily assess directional prediction, overlooking the structural complexity of derivative markets. Option trading introduces fundamentally different challen...
Hao-Chen Luo, Yi-Fan Li, Binh Minh An et al.· 0 citations
While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading...
Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial markets to content moderation to hiring. We show that improving individual model capability can degrade rather than improve system-level outcomes. We hypothesize that shared training and architectures can lead m...
Jillian Ross, Eric C. So, Zoe De Simone et al.· 0 citations
Investment institutions exert substantial influence on asset prices, liquidity, and market stability, making the ability to forecast their portfolio adjustments both academically and practically important. We study modeling investment institutions by fine-tuning large language models (LLMs) to predict next-quarter chan...
Yu-Xiang Cheng, Nan-Jiang Du, Qi-Da Mou et al.· IEEE Conference on Computati...· 0 citations
Large language models (LLMs) are increasingly used for financial decision-making, yet it remains unclear whether improvements in reasoning quality translate into better economic outcomes. We investigate this question using a multi-agent debate framework for portfolio allocation in historical market simulations, where s...
Ju-Li Huang, A. Alrassan, D. Harischandra et al.· 0 citations
Large language models (LLMs) have shown strong potential for financial analysis and trading, but direct trading remains challenging because the predictive capabilities required can vary across assets, decision fields, and market conditions. Existing LLM-based trading systems either coordinate human-defined external exp...
Chang Zhou, Xing-Tong Yu, Minbin Huang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.