Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning
Self-Verification via Reinforcement Learning (SVRL), an RL-only finetuning framework that trains multimodal agents to verify and filter retrieved evidence within their own reasoning traces, reducing reliance on external verifiers at inference time is presented.
Vishwas Sathish, Viresh Ranjan, Xin-Liang Zhu et al.
· 1 citation