Skip to content
Review

Confidence Estimation for Financial Vision-Language Models in Chart and Document Understanding

Aug 2026 · 0 citations · 32 references
Computer Science

TL;DR

This work evaluates seven confidence estimators, three inference-only and four trained internal probes, across five open-weight LVLMs and four conditions from three financial visual question-answering benchmarks, one bilingual; every probe is trained only on natural images and applied to finance without adaptation, so the results measure out-of-distribution transfer.

Abstract

LVLMs are increasingly used to read financial charts, tables, and documents, where a single misread figure can move a decision and the most authoritative-looking answer is sometimes one the model produced without reading the exhibit. The operational question is therefore trust, not accuracy: which answers can be acted on, and which escalated to a reviewer. We evaluate seven confidence estimators, three inference-only and four trained internal probes, across five open-weight LVLMs and four conditions from three financial visual question-answering benchmarks, one bilingual; every probe is trained only on natural images and applied to finance without adaptation, so the results measure out-of-distribution transfer. Three findings hold. First, the scarce property is calibration, not ranking: the inference baselines rank correct above incorrect answers competitively but are badly overconfident, calibration error far above what a threshold can tolerate, and only the trained probes produce a thresholdable score. Second, reliability is structured rather than global, along two axes a practitioner can read directly: the best estimator shifts with both model and task, none leading more than eight of twenty (model, condition) cells, and a controlled bilingual contrast exposes an apparent language robustness as a composition artifact that dissolves once models are read one at a time. Third, cast as deferral under an error budget, how much can be safely automated is set first by the model's competence and only narrowed by its confidence, so deferral clears a real share of the easiest condition and almost none of the hardest, near zero at a strict 5% budget. Two trained probes carry the calibration a deferral policy needs, and among them only the grounding-aware one lowers its confidence on answers a model gives without using the figure, separating detected non-grounding from a fluent guess.

View source

Similar papers

Review Aug 2026

Uncertainty-Aware Decision Making in Multimodal Large Language Models

This survey organizes the literature on uncertainty-aware MLLMs around a decision-centered framework: uncertainty sources give rise to observable signals, signals must be calibrated or controlled for risk, and calibrated uncertainty should determine the system action.

Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed · 0 citations
#machine learning Preprint Aug 2026

HalluPrism: When Multimodal Uncertainty Should Diagnose, Not Decide

HalluPrism, a behavioral diagnostic that re-runs an answer after visual degradation, blank-image replacement, and grounding or relation checks is proposed, a behavioral diagnostic that separates failure diagnosis from abstention scoring.

Aman Prakash, Sourish Dasgupta, Tanmoy Chakraborty · 0 citations
Review Aug 2026

Which Source Wins? Task-Dependent Reliance in Vision-Language Models

Modality reliance in VLMs is not fixed, but varies across tasks, evidence structures, models, and evaluation settings, which shows that modality reliance in VLMs is not fixed, but varies across tasks, evidence structures, and evaluation settings.

Rodela Ghosh, Aviral Gupta, Guang-Jing Wang · 0 citations
#natural language process... Preprint Sep 2026

Chronologic: Measuring Language Models'Ability to Represent the Past

Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct an...

Ted Underwood, Zi-Liang Qiu, Sarah Griebel et al. · 0 citations
#natural language process... Preprint Sep 2026

An Analysis of Training-Free Self-Reported Confidence in Language Models

Large language models can report a numerical confidence together with generated content, but it is unclear whether this report is more than calibrated rhetoric. We analyze three training-free signals: confidence verbalized with the answer, post-hoc $P(\mathrm{True})$, and agreement with three additional generations on...

L. Meyer, Sofia Rossi, Wei Chen et al. · 0 citations
Review Aug 2026

Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off

This survey traces attention from Bahdanau-Luong alignment through the Transformer and into vision architectures, and reviews fixed and learned sparse attention, linear attention, IO-aware exact algorithms including FlashAttention, and state-space alternatives including Mamba.

Aditya Singh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.