Text-to-image diffusion models enable data-efficient"mimicry"attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-energy watermarks, yet recent studies show these signatures are br...
Hao-Yang Li, Ruo-Xi Sun, Qing-Qing Ye et al.· 0 citations
Multimodal training relies on image-text corpora collected from external sources, creating opportunities for attackers to poison the data. Stealthy attacks can preserve plausible image-text pairs while concealing the differences used by detectors, so apparently clean data can still redirect the trained model. We theref...
Hao-Yang Li, Ya-Xin Xiao, Lin-Yan Dai et al.· 0 citations
Scalable oversight aims to verify the behaviour of agents whose capabilities exceed those of their overseers. AI debate has been proposed as an oversight solution in which competing agents help a resource-limited verifier assess claims that it cannot reliably evaluate unaided. Much of its promise rests on incentivizing...
Rayne Holland, Li-Ming Zhu, Jason Xue· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.