Skip to content

Author

Xiao-Mao Luo

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Unread or Unenforced? Separating Representation from Enforcement Failure in Content Guards

When an encoded attack passes a content guard, the guard either never represented the payload's harmful content or represented it and failed to act. End-to-end attack success rate reports one number for both, yet the two have opposite remedies: one is a representational limit that more safety training cannot reach, the...

Haoyu Zhang, Yi Feng, Shi-Bo Zheng et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Refusing Everything Looks Safe: Restoring the Benign Arm to Encoded-Prompt Evaluation

Encoded-prompt attacks are evaluated almost entirely on their harmful arm: a benchmark sends obfuscated harmful requests and reports how often the model complied, and shows that on two of the four models the loss is caused by the protocol rather than by the character transformation, and on a third by the characters.

Haoyu Zhang, Hao-Wen Xu, Xiao-Mao Luo et al. · 0 citations
Preprint Jul 2026

Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks

Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain...

Haoyu Zhang, Zhuo-Xiang Wang, Shi-Bo Zheng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.