Preprint
Aug 2026
Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO
Refusal-Calibrated Group Relative Policy Optimization (RC-GRPO) is proposed, a calibrated RL strategy that strengthens the refusal ability of MLLMs while preserving localization performance, and enforces"None" outputs in rollouts for valid advantage estimation on negative samples and applies a penalty to prevent over-refusal on positives, achieving a balanced trade-off between accuracy and reliability.
Xuzheng Yang, Jun Ling, Tao Huang et al.
· 1 citation