Skip to content

Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning

Sep 2026 · 1 citation · 59 references
Computer Science

TL;DR

Self-Verification via Reinforcement Learning (SVRL), an RL-only finetuning framework that trains multimodal agents to verify and filter retrieved evidence within their own reasoning traces, reducing reliance on external verifiers at inference time is presented.

Abstract

Reasoning agents increasingly rely on external tools such as web search to answer complex queries. Reinforcement learning (RL) finetuning algorithms such as GRPO have improved long-form reasoning in text-only language models, particularly for coding and mathematics. Reliable tool use in multimodal agents, however, remains challenging because models must interpret text and images while integrating noisy retrieved evidence, often under sparse outcome-level supervision without explicit verification signals. We present Self-Verification via Reinforcement Learning (SVRL), an RL-only finetuning framework that trains multimodal agents to verify and filter retrieved evidence within their own reasoning traces, reducing reliance on external verifiers at inference time. SVRL also introduces a search-aware penalty that discourages unnecessary tool calls and a query-diversity reward that encourages diverse, well-formed search queries, providing fine-grained feedback on when and what to search. Finetuning Qwen-2.5-VL-7B with SVRL on only 5{,}000 visual question answering examples yields consistent gains in multi-hop VQA generalization and tool efficiency across benchmarks. Overall, SVRL narrows the gap between compact agents and much larger proprietary models while requiring substantially lower training and inference cost.

View source

Similar papers

Preprint Aug 2026

StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning

This work introduces StructReward, a compute-efficient framework that provides dense reinforcement signals through structured step-level reward alignment and substantially reduces the computational overhead of multimodal reinforcement learning.

Yifan Li, Ruxi Sun, Tong-Zhou Zhao · 0 citations
#artificial intelligence Preprint Sep 2026

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

UnifiedPlayers, a cooperative framework comprising a Planning Player that generates tasks, an Execution Player that produces multi-turn trajectories with Python tool calls, and an Evaluation Player that constructs executable verifiers, highlights cooperation among specialized players as a promising path toward self-enh...

Wen-Jie Liao, Liang Zhao, Ze-Hong Cao · 0 citations
Preprint Aug 2026

Video-FLAIR: Not Whether to Reason, But How

Video-FLAIR is introduced, a training framework that learns to select the appropriate reasoning mode for each query using reinforcement learning, and yields a supervision signal for learning adaptive reasoning without per-query annotations.

Yogesh Kulkarni, Pooyan Fazli · 0 citations
#artificial intelligence Preprint Aug 2026

AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing

This work proposes AgenticRag-R1, a RL framework that deeply integrates reasoning, retrieval, and memory via a memory stack and fine-grained action space, supported by hierarchical action-aware rewards and an information-aware trajectory rejection strategy to enable effective long-horizon learning.

Xin-Ke Jiang, Yue Fang, Zhi-Bang Yang et al. · 2 citations
#artificial intelligence Preprint Aug 2026

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

This work introduces Co-RL, a framework in which multiple decoupled models, sharing no parameters, are simultaneously optimized through RL using rewards derived from their peers, and shows that unsupervised reasoning can emerge through cooperative multi-agent training.

Yunhao Yang, Yuexin Bian, Yun-Jie Tian et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.