Skip to content
Open access

Boundary Sampling for Efficient Model Extraction

Doron Ben Chayim Maor Biton Dor Eyal Lenga Yisroel Mirsky
Jul 2026 · ACM Transactions on Intelligent Systems and Technology · Vol 17, pp. 1-35 · 0 citations · 49 references

TL;DR

This work proposes a novel data-free model extraction attack that substantially outperforms current methods in efficiency, accuracy, and overall effectiveness and offers extensions to the algorithm to enable it to work on complex models.

Abstract

This work proposes a novel data-free model extraction attack that substantially outperforms current methods in efficiency, accuracy, and overall effectiveness. Conventional black-box attacks depend heavily on treating the victim model as an oracle to label a large number of samples, primarily within high-confidence regions. This strategy not only demands an excessive number of queries but also often leads to the extraction of models with lower accuracy and limited transferability. In contrast, our method shifts focus to sampling low-confidence regions (along the decision boundaries) and leverages an evolutionary algorithm to enhance the sampling process. This approach dramatically reduces the query requirement by a factor of 10x to 600x, while also increasing the accuracy of the extracted model. Furthermore, our method achieves improved boundary alignment, significantly enhancing the transferability of adversarial examples from the extracted model to the victim, increasing the attack success rate from an average of 60% to 82%. Remarkably, these improvements are accomplished under a strict black-box scenario with soft-label (class-probability) query access, and no prior knowledge of the target model’s architecture or data distribution. Finally, we offer extensions to the algorithm to enable it to work on complex models: with high resolution, many classes, and even models with class imbalance such as anomaly detectors. Our attack is thoroughly evaluated on multiple image datasets with varying resolutions and is benchmarked against many state-of-the-art model extraction techniques. Additionally, to illustrate the versatility and robustness of our method, we conduct extensive experiments on four tabular datasets that vary in class numbers and sizes.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

No Free Efficiency: Revisiting the Trade-off Between Training Efficiency and Model Vulnerability

Across vision and language models, it is shown that efficiency-oriented training increases susceptibility to adversarial and privacy attacks, and is called for a paradigm shift toward multi-objective training that jointly optimizes for performance, cost, and security.

Yi-Yong Liu, Jun Sakuma, Michael Backes et al. · 0 citations
Sep 2026

Towards Transferable Black-Box Attack via Minimizing Maximum Model Discrepancy.

Adversarial attacks on black-box models, which operate without direct access to the target system, present a significant challenge due to the lack of a foundational theory for the transferability of adversarial examples. This paper introduces a paradigm shift in black-box adversarial attacks by minimizing model discrep...

An-Qi Zhao, Tong Chu, Ya-Hao Liu et al. · 0 citations
Open access Sep 2025

Text Adversarial Attacks With Dynamic Outputs

Text adversarial attack methods are typically designed for static scenarios with fixed numbers of output labels and a predefined label space, relying on extensive querying of the victim model (query-based attacks) or the surrogate model (transfer-based attacks). However, real-world applications often involve non-static...

Wen-Qiang Wang, Si-Yuan Liang, Yangshijie Zhang et al. · 2 citations
Preprint Aug 2026

Diversity Matters: Distributional Feature Coverage Sample Selection for Data-Efficient Backdoor Attacks

DFCS is proposed, a training-free, trigger-agnostic method that clusters fixed pretrained features into one region per poisoning slot and selects the centroid-nearest sample from each region and supports distributional feature coverage as an effective selection principle for low-budget dirty-label backdoor attacks.

Yi Yang, Xiaoke Chen, Jin-Yang Huang et al. · 0 citations
Preprint Aug 2026

A Convolutional Layer Activation Dimensionality Reduction for Out-of-Distribution and Adversarial Attack Detection Methods

This paper carefully analyzes two state-of-the-art detection methods and their dimensionality reductions for convolutional layers and develops a novel reduction method with a controllable high-compression level.

Leandro de Souza Rosa, Lorenzo Capelli, Clara Nunes Barrancos et al. · 0 citations
Conference Sep 2026

Enhancing adversarial defense robustness through sensitive prediction region mining

Current adversarial defense methods often rely on specific perturbation generation techniques, which face challenges such as limited generalization performance and high computational costs. This paper addresses these issues by examining the response characteristics of intelligent learning models to sensitive adversaria...

Qian Li, Di Wu, Saiyu Qi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.