Text Adversarial Attacks With Dynamic Outputs
Abstract
Text adversarial attack methods are typically designed for static scenarios with fixed numbers of output labels and a predefined label space, relying on extensive querying of the victim model (query-based attacks) or the surrogate model (transfer-based attacks). However, real-world applications often involve non-static outputs. Large Language Models (LLMs), for instance, may generate labels beyond a predefined label space, and in multi-label classification tasks, the number of predicted labels can vary dynamically with the input. We refer to these two forms of label variability, namely dynamic label content and dynamic label number, as the Dynamic Outputs (DO) scenario. Existing adversarial attack methods are not designed for such settings and are therefore not directly applicable under DO conditions. To address this gap, we introduce the Textual Dynamic Outputs Attack (TDOA) method, which employs a clustering-based surrogate model training approach to approximate dynamic fine-grained outputs using static coarse-grained labels, thereby constructing a conventional single-output surrogate attack space. This approximation is expected to be effective when the dynamic outputs exhibit coherent semantic structures and the generated coarse-grained clusters are sufficiently reliable. To improve attack effectiveness, we propose the farthest-label targeted attack strategy, which guides perturbations toward the semantically most distant coarse-grained label, aiming to induce larger changes in the victim model’s outputs. We extensively evaluate TDOA on five datasets and ten victim models (e.g., GPT-4o, GPT-4.1), showing its effectiveness in crafting adversarial examples and its strong potential to compromise LLMs with limited access. With five queries per text, TDOA achieves a maximum attack success rate of 80.8%. Additionally, we find that TDOA also achieves SOTA performance in conventional static output scenarios, reaching a maximum ASR of 82.7%. We further extend TDOA from dynamic-label settings to open-ended generative outputs through machine translation experiments, rather than treating translation as conventional classification with an unbounded label space. In this extended setting, TDOA improves the previous best results by up to 0.11 in both RDBLEU and RDchrF.