Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language interaction, yet their safety alignment remains vulnerable to jailbreak attacks. A key challenge is that safety behavior learned in the textual space does not reliably transfer to fused cross-modal representations, leaving mul...