Jul 2026· Review of Behavioral Economics· 0 citations· 24 references
TL;DR
A replicable methodology is introduced, findings across two architecturally distinct LLMs from different developers are extended, and it is demonstrated that deliberate prompt design meaningfully reduces AI decision bias.
Abstract
This study investigates behavioral biases of generative artificial intelligence (AI) models, specifically GPT-4o and Claude-Haiku-4.5, in inventory management using the newsvendor problem. This study compares AI decision-making with human-subject experiments to assess whether large language models (LLMs) replicate human cognitive bias and to identify prompt-design strategies that improve alignment with optimal outcomes.
Controlled newsvendor experiments were conducted with generative AI models, mirroring established human-subject laboratory protocols. Prompt framing was systematically varied across three modifications: removing explicit waste and missed-profit information, simplifying instruction format and providing explicit optimization formulas. Results were benchmarked against normative economic predictions and existing human behavioral findings.
Generative AI exhibits human-like human biases including risk aversion, loss aversion and demand chasing, but exhibits a stronger demand-chasing tendency than human participants. It responds to hypothetical incentives and displays bounded rationality. Prompt design significantly influences decision quality, producing decisions closer to theoretical benchmarks.
This study empirically tests generative AI behavioral biases within a structured operations management experiment. It introduces a replicable methodology, extends findings across two architecturally distinct LLMs from different developers, and demonstrates that deliberate prompt design meaningfully reduces AI decision bias. The study also contributes a conceptual distinction between functionally analogous behavioral patterns and intrinsic psychological dispositions in LLMs, offering a more precise interpretive framework for AI decision-making research in operational contexts.
Benevolence bias is identified and measure, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions, and is easy to diagnose and straightforward to fix.
Yuanzi Li, Jun-Hao Wang, Minghui Liu et al.· 0 citations
Human choice behavior, including route choice, exhibits systematic behavioral biases that deviate from the assumptions of full rationality. Cumulative prospect theory (CPT) has been widely recognized as an effective framework for characterizing such behavioral patterns. However, its large-scale application, particularly in simulation and agent-based modeling, critically depends on specifying individual-level CPT parameters, which remain a major bottleneck. Conventional approaches typically rely on surveys and controlled experiments to calibrate CPT parameters, yet these methods are difficult to generalize and often fail to capture the full diversity of human decision-making. To address this challenge, this paper investigates whether large language models (LLMs) can reproduce human behavioral biases in choice-making without explicit specification of prospect-theoretic parameters. Using route choice as a representative scenario, we design a behavioral evaluation framework and systematically compare LLM-generated decisions with established human behavioral patterns predicted by CPT. Experimental results demonstrate that LLMs are capable of reproducing non-rational human choice biases and can exhibit decision behaviors consistent with prospect-theoretic effects under uncertainty. These findings suggest that generative AI models may provide a scalable alternative for modeling human decision processes and offer a promising foundation for next-generation large-scale agent-based simulation and AI-driven behavioral research.
Jiangtao Han, Shoufeng Ma, Shuxian Xu et al.· 0 citations
Applying the method to 119,147 decisions made by 78,657 subjects from more than 35 countries across 10 classic economic game roles, it is found that human behavior can be closely matched using three dimensions: Risk Aversion, Strategic Sophistication, and Trust.
Matthew O. Jackson, Benjamin S. Manning, Yutong Xie et al.· 0 citations
LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief updates, comparing LLM outputs 1-to-1 against ground truth data from 391 UK participants on Prolific, who updated their stances on three discussion topics after reading Reddit comments. Each participant was simulated by an LLM conditioned on a persona derived from their demographic and personality trait data. We find that some LLMs (Qwen3-32B and GPT-5-Mini) can match the human post-stance distribution, but only when given participants'actual initial stances. All six models fail to simulate initial stances themselves and to produce faithful belief updates from self-generated stances. Three systematic biases emerge across all models: overrepresentation of neutral positions, more frequent but smaller belief shifts than humans, and a failure to rank comments by convincingness. Demographic and personality trait personas had no consistent effect on fidelity. LLM simulations of human belief dynamics are only reliable when grounded in realistic starting conditions, that current multi-round social media simulations rarely provide.
Sebastian Pohl, Harsh Mehta, Pranav Mambayil et al.· 0 citations
Repeated human-AI interaction is often analyzed through pooled belief-updating slopes: users observe AI successes and failures, revise reported beliefs in the feedback-consistent direction, but appear conservative on average. We show that such averages can obscure an important distinction between whether an elicited belief report changes at all and how it changes conditional on movement. We refer to this measurement-aware decomposition as the belief update gate. Reanalyzing a multi-task human-AI decision-making dataset with 240 participants, 7,200 trials, and three task domains, we find substantial non-movement in reported beliefs: 67.3% of trial-level belief changes are exactly zero, and 76.4% are smaller than five percentage points. Separating non-moving from moving reports changes the descriptive interpretation of pooled conservatism: the within-trajectory slope rises from 0.494 overall to 0.949 among rows with nonzero movement. Since this latter estimate conditions on observed movement, we interpret it as a descriptive decomposition rather than as evidence of a near-Bayesian latent learning process. Complementary hurdle style analyses (i.e., modeling zero vs. non-zero changes before predicting update magnitude) show that the absolute discrepancy between feedback and entering belief predicts whether a report changes, while the signed feedback discrepancy predicts the direction and magnitude of change among reports that move. Importantly, observed non-movement does not distinguish genuine latent belief inertia from small unexpressed updates, rounding, or other reporting processes. These findings show that calibration analyses of repeated human--AI interaction should distinguish visible non-movement in elicited belief reports from updating conditional on movement rather than treating reported beliefs as a single continuous updating process.
Shreyan Biswas, Alexander Erlei, U. Gadiraju· 0 citations