SAVER (Selective Auditing of Verbal Evidence for Error Recovery), a lightweight, rule-based method that parses VLM responses for this evidence and triggers structured reprompting only when evidence is missing or inconsistent, significantly improves accuracy on tasks where errors stem from the model failing to articulate what it saw.
Abstract
Vision-language models (VLMs) frequently fail at visual change reasoning, even when their vision encoders contain sufficient information. We observe that correct VLM outputs tend to contain explicit verbal evidence (object names, colors, spatial locations) that supports the claimed change, while incorrect outputs often lack such evidence. We propose SAVER (Selective Auditing of Verbal Evidence for Error Recovery), a lightweight, rule-based method that parses VLM responses for this evidence and triggers structured reprompting only when evidence is missing or inconsistent. Across three change detection benchmarks and four VLMs, SAVER significantly improves accuracy on tasks where errors stem from the model failing to articulate what it saw (expression failures), with gains up to +25.8% on CLEVR-Change. The evidence patterns can also be generated by an LLM in a single call, matching the hand-tuned gate on CLEVR-Change. Ablation experiments confirm that the evidence gate, not reprompting alone, drives the improvement.
Results show that process priors are most useful when aligned with the scene's specific decision boundary, motivating boundary-aware prior selection for process-grounded visual reasoning.
Yu-Yang Dai, Bo-Fei Huang, Hong-Bo Zhang et al.· 0 citations
Locator-Critic (LOCI) is proposed, a training-free framework that decouples visual search from evidence verification and improves accuracy for both open-weight models like Qwen3-VL and proprietary models like Gemini 2.5 Pro.
Walid Bousselham, Mathilde Caron, Arsha Nagrani et al.· 0 citations
Findings show that configuration changes can affect how models use information they can still read, and attention interventions in LLaVA-NeXT suggest that configuration changes can weaken the use of readable information during answering.
Ding-Yang Lin, Ying-Feng Luo, Cheng-Long Wang et al.· 0 citations
This work introduces FailBench, a benchmark for robot failure detection comprising 2,197 manipulation attempts across 14 public sources (12 real-world, 2 simulated), and shows that input-level intervention--spatially localizing and cropping outcome-relevant regions--improves the top detector by 2.4 percentage points wi...
The analysis of the imitation objective shows how narrow conditional action support can leave grounded and instruction-keyed solutions indistinguishable on the demonstrations, and motivates Equivariant Counterfactual Training (ECT), which acts at two levels.
Hung-Jen Chen, Yue-Ling Hou, Yan-Hong Chen et al.· 0 citations
Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterf...