Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails
This work studies the problem in web-agent guardrails, where a VLM judges whether on-screen text conflicts with a user instruction, and proposes two training-free interventions that improve grounding without changing verdicts.
Suyoung Lee, Myungsub Choi
· 0 citations