Skip to content
Preprint

Steal the Patch Size: Adversarially Manipulate Vision-Language Models

Jun 2026 · 0 citations · 24 references
Computer Science

TL;DR

A black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models, including the visual patch size and input preprocessing pipeline, and it is shown that such leakage enables preprocessing-aware transfer attacks and model-targeted adversarial manipulation.

Abstract

We present a black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models (VLMs), including the visual patch size and input preprocessing pipeline. The key idea is a task-level side channel induced by ViT-style patchification: when a synthetic grid image is aligned with the hidden patch grid, boundary cues are erased at tokenization, causing periodic accuracy drop. By sweeping the grid cell size and measuring these collapses, we infer the patch size; by introducing padding and a consistency-check test, we further identify whether preprocessing is dynamic- or fixed-resolution and recover the target resize resolution. Across open-source Qwen-VL variants and proprietary models including GPT and Claude, we reliably recover tokenizer-related parameters. Finally, we show that such leakage enables preprocessing-aware transfer attacks and model-targeted adversarial manipulation.

View source

Similar papers

Preprint Aug 2026

Adversarial Attacks on Deep OCR Systems

Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at low token cost. However, its increased complexity may introduce new security vulnerabilities. In this paper, we present, to the best of our knowledge, the first pure black-box adversarial attack against a generative OCR vision-language model, where only the decoded string can be queried and no gradients, logits, or model internals are available. We recast the attack as a zeroth-order optimization problem driven by a bounded scalar loss defined directly on the string output via sequence similarity, and estimate the gradient with a random-direction finite-difference scheme whose query cost is independent of the image dimension. An Adam update with ell_infinity projection yields imperceptible perturbations for both untargeted and targeted objectives. Pilot experiments on Deep-OCR validate the string-only attack and evaluation pipeline and expose severe qualitative decoder failures, including repetition, truncation, and prompt leakage. They also show that controlled targeted rewriting remains substantially harder than untargeted degradation; we avoid claiming targeted success until the pre-registered evaluation is complete.

Wenbo Sun, Hong-Zong Li, Yanyun Wang et al. · 0 citations
Preprint Jul 2026

GhostPrompt: Cross-Image Adversarial Prompt for Vision-Language Models

Vision-Language Models (VLMs) are known to be vulnerable to adversarial attacks, where subtle perturbations to images or texts induce erroneous outputs. However, most text-based attacks are adapted from language-model-centric methods, in which the visual input is fixed during optimization, resulting in adversarial prompts that are tied to specific images and thus limiting their attack effectiveness. To this end, we first introduce a new research perspective: cross-image transferability for adversarial prompts. We then propose GhostPrompt, an adversarial prompt that is optimized once and reused to steer VLM outputs toward attacker-specified responses across diverse images. GhostPrompt employs a joint optimization that distills image-invariant adversarial features into the prompt by"worst-case"generation. Specifically, it alternates between constructing hard visual conditions for the current prompt and updating the prompt to remain effective under these conditions. Extensive experiments on prevalent VLMs verify that \ourmethod achieves an improvement of over 30% in attack success rates compared to state-of-the-art (SoTA) baselines, while reducing computation time by ~70%. Our code is avalable at https://github.com/Ye-ze-yu/GhostPrompt.

Li Zeng, Ze Ye, Meng Xie et al. · 0 citations

A noise-based defense for stealthy backdoor attacks in large vision-language models

Large vision-language models rely on pretrained vision encoders to translate images intofeature representations used by downstream language models. This creates a security riskwhen the encoder is compromised by a stealthy backdoor attack, such as BadVision, where asubtle trigger causes an image to be mapped toward an attacker-chosen target representationwhile clean inputs remain largely unaffected. Because the model behaves normally understandard evaluation, these attacks are difficult to detect.This thesis investigates controlled noise injection as a lightweight input-side defenseagainst BadVision-style backdoors. The proposed approach adds small perturbations toinput images before they enter the vision encoder, with the goal of disrupting the triggerwhile preserving the semantic content of clean images. Several perturbation types are evaluated, including Gaussian noise, random noise, salt-and-pepper noise, low-frequency noise,geometric transformations, occlusion, scaling, rotation, and channel-based distributions.Experimental results show that geometric and channel-based transformations have limitedeffect on the backdoor, while pixel-level statistical perturbations significantly reduce targetsimilarity, increase feature-space distance from the attacker’s target representation, and lowerattack success. These findings suggest that stealthy encoder-level triggers depend on fragilestatistical patterns and can be weakened through controlled noise injection without requiringretraining of the full multimodal model.

James Patrick Donohue · 0 citations
Preprint Jul 2026

Lights, Camera, Malfunction: When Illumination Robustness Leaves VLA Models Blind to Color

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for general-purpose robot manipulation; however, their transition to real-world environments reveals vulnerabilities to minor environmental perturbations. We propose FLARE, an optimized physical spotlight attack framework that exploits these vulnerabilities via targeted illuminations, dropping baseline task success rates to zero without any access to model internals. While adversarial training is the standard countermeasure, we identify a critical and previously underestimated defensive pitfall: naive data augmentations incorrectly condition VLA models to discard color as noise, collapsing their visual perception into a purely shape-biased processor. We expose this degradation through a diagnostic grayscale evaluation, in which the defended model maintains high success rates on grayscale inputs, while its success rate on benign, color-dependent real-world tasks drops to at most 47.5%, well below the undefended baseline. To address this, we propose ChromaGuard, a chroma-preserving adversarial training method. On a physical 6-DoF robotic platform, we demonstrate that ChromaGuard achieves 97.5% and 92.5% success rates in benign and attacked color-dependent tasks, respectively.

Marino Watanabe, Takami Sato, Kentaro Yoshioka · 0 citations
Preprint Jul 2026

A New Kind of Adversarial Example: Measuring the Human-Model Gap, and Its Relationship to OOD Detection

Almost all adversarial attacks add an imperceptible perturbation to fool a model. We instead study the opposite: a large, clearly visible perturbation that causes the model to keep its original, correct prediction, even though a human would no longer recognize the image. Prior work showed such examples can be generated at scale but left three questions untested: whether humans really perform worse than the model, whether standard out-of-distribution (OOD) detection and calibration tools catch it, and whether existing defenses mitigate it. We answer all three on MNIST, CIFAR-10, and ImageNet. (i) An independent recognizer proxy drops to ~49% on CIFAR-10 while the model stays at 100% -- a gap a small human pilot (N=5) corroborates directly and that is not explained by signal loss (a matched-magnitude Gaussian control degrades recognizability faster); a CLIP zero-shot proxy confirms the gap at ImageNet scale too. (ii) Confidence- and energy-based OOD detectors and calibration are structurally blind (0% detection, ECE ~= 0), while a feature-space Mahalanobis detector flags 100% -- but is evaded by an adaptive attacker at no cost to success. (iii) No classical defense, including adversarial training (45% robust accuracy), reduces attack success (correlation with large-epsilon_l resistance r ~= 0). A mechanistic analysis further shows the attack destroys low-level texture far faster than edge/shape structure.

A. Borji · 0 citations
Open access Jul 2026

On Success and Simplicity: A Second Look at Transferable Vision–Language Attack Pipeline

This paper identifies three previously overlooked issues caused by inappropriate cross-modal interactions and excessive operations in the Simple Vision-Language Attack (SimVLA) pipeline, and proposes the SimVLA, which observably improves transferability and efficiency.

Yuchen Ren, Zhengyu Zhao, Chenhao Lin et al. · 0 citations