Skip to content
Preprint

SAVER: Selective Auditing of Verbal Evidence for Error Recovery in VLM Change Reasoning

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

SAVER (Selective Auditing of Verbal Evidence for Error Recovery), a lightweight, rule-based method that parses VLM responses for this evidence and triggers structured reprompting only when evidence is missing or inconsistent, significantly improves accuracy on tasks where errors stem from the model failing to articulate what it saw.

Abstract

Vision-language models (VLMs) frequently fail at visual change reasoning, even when their vision encoders contain sufficient information. We observe that correct VLM outputs tend to contain explicit verbal evidence (object names, colors, spatial locations) that supports the claimed change, while incorrect outputs often lack such evidence. We propose SAVER (Selective Auditing of Verbal Evidence for Error Recovery), a lightweight, rule-based method that parses VLM responses for this evidence and triggers structured reprompting only when evidence is missing or inconsistent. Across three change detection benchmarks and four VLMs, SAVER significantly improves accuracy on tasks where errors stem from the model failing to articulate what it saw (expression failures), with gains up to +25.8% on CLEVR-Change. The evidence patterns can also be generated by an LLM in a single call, matching the hand-tuned gate on CLEVR-Change. Ablation experiments confirm that the evidence gate, not reprompting alone, drives the improvement.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

LOCI: A Locator-Critic with Refinement Loop

Locator-Critic (LOCI) is proposed, a training-free framework that decouples visual search from evidence verification and improves accuracy for both open-weight models like Qwen3-VL and proprietary models like Gemini 2.5 Pro.

Walid Bousselham, Mathilde Caron, Arsha Nagrani et al. · 0 citations
#small language model Preprint Sep 2026

Reading Right, Answering Wrong: How Visual Configuration Changes Affect Evidence Use in VLMs

Findings show that configuration changes can affect how models use information they can still read, and attention interventions in LLaVA-NeXT suggest that configuration changes can weaken the use of readable information during answering.

Ding-Yang Lin, Ying-Feng Luo, Cheng-Long Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FailBench: How Reliable are VLMs at Judging Robot Task Success?

This work introduces FailBench, a benchmark for robot failure detection comprising 2,197 manipulation attempts across 14 public sources (12 real-world, 2 simulated), and shows that input-level intervention--spatially localizing and cropping outcome-relevant regions--improves the top detector by 2.4 percentage points wi...

Zaruhi Navasardyan, Tatul Danielyan, Hrant Davtyan · 0 citations
#machine learning Preprint Sep 2026

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

The analysis of the imitation objective shows how narrow conditional action support can leave grounded and instruction-keyed solutions indistinguishable on the demonstrations, and motivates Equivariant Counterfactual Training (ECT), which acts at two levels.

Hung-Jen Chen, Yue-Ling Hou, Yan-Hong Chen et al. · 0 citations
#machine learning Preprint Sep 2026

From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models

Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterf...

Rong-Yu Xu, Prayag Tiwari, Shao-Lei Zhang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.