Multiscale Frame Transformation for Transferable Adversarial Attacks on Remote Sensing Scene Classification
Abstract
Transferable adversarial attacks provide an important means of evaluating the black-box security of remote sensing scene classification models. However, the existing spatial input transformations commonly rely on predefined block shapes and limited partition granularities, providing insufficient coverage of the diverse orientations and spatial scales in remote sensing scenes. We propose a training-free multiscale frame transformation (MFT) method that samples orientation-complementary frame patterns at different granularities and constructs randomized spatial views through frame shuffling and lightweight intraframe transformations. During each attack iteration, gradients from multiple MFT views are aggregated to reduce sensitivity to individual spatial layouts. Experiments on NWPU-RESISC45 and AID show that the MFT achieves the highest average black-box attack success rate (ASR) in all four dataset–surrogate settings and transfers effectively across convolutional neural network (CNN), transformer, and contrastive language–image pretraining (CLIP) architectures. Across 32 black-box target settings, the MFT achieves an overall average ASR of 91.64%, exceeding the best competing baseline by 2.66 percentage points. Ablation studies further support the benefits of orientation complementarity and cross-granularity integration.