Encoded-prompt attacks are evaluated almost entirely on their harmful arm: a benchmark sends obfuscated harmful requests and reports how often the model complied, and shows that on two of the four models the loss is caused by the protocol rather than by the character transformation, and on a third by the characters.
Haoyu Zhang, Hao-Wen Xu, Xiao-Mao Luo et al.· 0 citations
An empirical safety– utility ceiling for the non-iterative recovery-based defenses the authors evaluate, recurring across every guard and both target VLMs, is exposed.
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain...
Haoyu Zhang, Zhuo-Xiang Wang, Shi-Bo Zheng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.