Preprint
Jul 2026
Open Models, Open Risks: Measuring Unsafe Generation in Text-to-Image Models In the Wild
A large-scale empirical study of in-the-wild T2I safety through the lens of jailbreak, showing that detector-only jailbreak metrics substantially overestimate practical risk over in the wild due to semantic drift and generation artifacts and introducing Advanced ASR to better capture semantically valid and visually plausible unsafe generation.
Peilin Han, Yang Liu, Yilong Yang et al.
· 0 citations