Skip to content
Preprint

Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows

Aug 2026 · 0 citations
Computer Science

TL;DR

A retrieval-integration gap in long-context financial analysis is identified and it is found that a risk disclosure's influence on investment judgments falls to the experimental noise floor even as direct retrieval remains accurate.

Abstract

Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they can retrieve, not whether retrieved information affects their judgments. We identify a retrieval-integration gap in long-context financial analysis. Holding focal-firm information fixed and varying only unrelated context from 2,000 to 128,000 tokens, we find that a risk disclosure's influence on investment judgments falls to the experimental noise floor even as direct retrieval remains accurate. The pattern replicates across model families and judgment tasks and in experiments removing real disclosures from actual 10-K filings. More capable models postpone but do not eliminate the gap. Causal memory interventions show that compressed summaries and source-text lookup jointly transmit disclosures into judgments. Workflow architecture determines whether this transmission succeeds: chunk-and-summarize pipelines evict relevant information, whereas a targeted, structured restatement adjacent to the decision restores its influence. AI analyst performance is therefore jointly determined by model capability and workflow architecture. Retrieval-based evaluations can certify systems whose investment judgments ignore information they demonstrably retrieved.

View source

Similar papers

Review Sep 2026

Mining Meaning: Measurement Error in AI-Assisted Literature Reviews

Researchers increasingly use generative AI, particularly large language models (LLMs), to automate tasks across the research pipeline. We study the reliability of these tools at the reading, classification, and synthesis of large bodies of academic literature. We frame LLM-assisted literature reviews as a measurement p...

J. Michler, Kiera Douglas, Anna Josephson · 0 citations
#artificial intelligence Preprint Aug 2026

FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems

Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed across invoices, purchase orders, approvals, allocations, payments, l...

Prof. S. B. Ghawate · 2 citations
#artificial intelligence Preprint Sep 2026

Does AI Assistance Leave a Temporal Fingerprint? Detecting Overreliance in AI-Assisted Writing and Programming

This work analyzes three public corpora: CoAuthor (1,447 keystroke-level co-writing sessions), RealHumanEval (editor telemetry from 243 programmer records), and a pre-LLM CS1 corpus as a human-only baseline, comparing minimal-AI work, collaborative AI use, and simulated wholesale delegation.

Eduardo Davalos, Yi-Ke Zhang · 0 citations
#machine learning Preprint Sep 2026

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...

Himil Vasava, Ming-Zhou Jiang · 0 citations
#natural language process... Preprint Sep 2026

FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability

FinFIRST is the first financial benchmark to jointly evaluate answers and supporting evidence through atomic rubrics, retaining final-answer correctness as the primary objective while making the supporting research process measurable, verifiable, and diagnosable.

Wen-Qing Wang, Hai-Tao Xiang, Xin-Yi Zhao et al. · 0 citations
Preprint Aug 2026

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

It is shown that prior scores, even when included only as context metadata, anchor judgments and systematically shift ratings toward their values, and effective mitigation must be validated for the intended model and task or domain.

A. Kapetanović, Kemal Altwlkany, Andro Merćep et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.