Conference
Open access
2026
B-APO: Bias-Targeted Adversarial Preference Optimization for Debiasing Multimodal Large Language Models
This work proposes B-APO (Bias-Targeted Adversarial Preference Optimization), which casts debiasing as a bias-targeted min-max game: it generates hard negatives by applying small adversarial perturbations in the latent space to maximally induce language-vision-prior reliance, and then performs preference alignment to enlarge the margin between clean and adversarial responses.
Pinlong Zhao, Zike Ding, Zengshu Ye et al.
· Annual Meeting of the Associ... · 0 citations