CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition
Text-to-image (T2I) models have substantially improved in language understanding, in-image text rendering, and visual composition, while their safety mechanisms do not always keep pace with these capabilities. This creates a cross-modal attack surface in which harmful semantics can remain inconspicuous in a serialized...