Preprint
Jun 2026
Steal the Patch Size: Adversarially Manipulate Vision-Language Models
A black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models, including the visual patch size and input preprocessing pipeline, and it is shown that such leakage enables preprocessing-aware transfer attacks and model-targeted adversarial manipulation.
Kai Hu, Akash Bharadwaj, Weicheng Yu et al.
· 0 citations