Findings show that financial-advice evaluations are shaped jointly by displayed attribution and message-level communication cues, which position disclosure not as a neutral transparency mechanism, but as an interpretive frame whose accuracy and interaction with message cues can shape trust and reliance.
Abstract
As generative AI increasingly becomes a common source of daily decision-making, including financial choices, it is critical to understand how people evaluate AI-generated financial advice. We conducted a preregistered vignette experiment (N = 285) in which substantive financial content---including facts, numerical values, recommendation direction, and core reasoning---was held constant while communication style varied across AI Financial Assistant (AI), Certified Financial Planner (Expert), and Online Community Forum (OC) advice. Displayed source attribution was independently manipulated through correctly labeled, unlabeled, and mislabeled conditions, allowing us to separate attribution effects from source-specific communication cues. Expert advice was rated more favorably than AI advice on 9 of 10 outcomes (|d|=0.20--0.47), and this advantage remained visible without source labels, where Expert advice outperformed AI advice on 8 of 10 outcomes (up to d=0.60). Correct labels added limited differentiation, whereas mislabeling increased ratings of AI advice for situational fit and overall quality (d=0.42 for each) and attenuated the Expert advantage in situational fit (d=-0.36). Descriptive analyses further showed that AI advice was most responsive to displayed attribution and, conversely, that advice-style differences were most visible under an AI label. These findings show that financial-advice evaluations are shaped jointly by displayed attribution and message-level communication cues. We position disclosure not as a neutral transparency mechanism, but as an interpretive frame whose accuracy and interaction with message cues can shape trust and reliance.
Knowing when to say"I don't know"is fundamental to human judgment, yet AI assistants offer a fluent answer to almost any question. In five experiments (N = 3,132; four preregistered, one direct replication), participants answered difficult questions and could always decline to respond. We engineered the questions so that AI advice was wrong, separating AI use from its accuracy. Merely having access to AI nearly eliminated participants'willingness to suspend judgment, and this held whether the advice was actively requested or simply displayed. Consequently, participants answered more questions but were correct about a third as often as when AI was unavailable-yet their confidence nearly doubled. Incentivizing accuracy and penalizing inaccuracy led participants to seek and follow AI advice less, answer more accurately, and suspend judgment more often, though still far less than when AI was unavailable. As AI suggestions grow ubiquitous and unsolicited, they may not simply affect answer accuracy; they may even alter the metacognitive threshold at which people decide whether they know enough to answer.
Chiara Marcoccia, Walter Quattrociocchi, Valerio Capraro· 2 citations
A replicable methodology is introduced, findings across two architecturally distinct LLMs from different developers are extended, and it is demonstrated that deliberate prompt design meaningfully reduces AI decision bias.
Jing-Jie Su, Yan Lang, Kay-Yut Chen· Review of Behavioral Economi...· 0 citations
It is indicated that AI-powered personal finance applications are consistently associated with improved financial confidence and more frequent, though not necessarily more diversified, investment activity among young users, with effects moderated by financial literacy, trust in automation, and platform design quality.
Shashank Adagond, D. R. G KARGAL· International Scientific Jou...· 0 citations
Applying the method to 119,147 decisions made by 78,657 subjects from more than 35 countries across 10 classic economic game roles, it is found that human behavior can be closely matched using three dimensions: Risk Aversion, Strategic Sophistication, and Trust.
Matthew O. Jackson, Benjamin S. Manning, Yutong Xie et al.· 0 citations
It is demonstrated that GRPO with a finance-grounded reward signal can produce substantially more useful business recommendations than commercial LLMs, and that a judge-independent causal audit is a valuable complement to, rather than a confirmation of, LLM-as-a-judge assessment in financial NLP.
Ofir Ben Shoham, Shrutendra Harsola, Vignesh T. Subrahmaniam et al.· 0 citations
A prespecified randomized algorithm audit of what causally moves large language model (LLM) assistants' recommendations, finding that gender and ethnicity were signaled through names following correspondence-audit methodology.
Syeda Anshrah Gillani, Mirza Samad Ahmed Baig· 0 citations