Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Aug 2026

Refusing Everything Looks Safe: Restoring the Benign Arm to Encoded-Prompt Evaluation

Encoded-prompt attacks are evaluated almost entirely on their harmful arm: a benchmark sends obfuscated harmful requests and reports how often the model complied, and shows that on two of the four models the loss is caused by the protocol rather than by the character transformation, and on a third by the characters.

Haoyu Zhang, Hao-Wen Xu, Xiao-Mao Luo et al. · 0 citations
#artificial intelligence Preprint Aug 2026

The Uncontrolled Variable: Vision-Language Refusal Is Conditioned on the Image-Attachment Interface, and Not Robust to Irrelevant Image Properties

We show that aligned vision-language models also condition refusal on a property of a request's form: whether an image is attached, holding everything the request asks fixed. Attaching a blank canvas, an image that cannot be read, cannot relate to the request, and is byte-identical across every prompt in its condition,...

Haoyu Zhang, Yi Feng, Han-Wen Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.