Jul 2026· International Journal of Innovative Science and Research Technology· 0 citations· 9 references
TL;DR
This survey provides a comprehensive synthesis of adversarial attacks and defense mechanisms in modern AI security, introducing a structured taxonomy categorizing attacks into evasion, poisoning, and model inversion strategies, evaluated across varying levels of attacker knowledge.
Abstract
The growing deployment of deep learning models in safety-critical domains has exposed the artificial intelligence
landscape to a widening array of adversarial threats, where imperceptible input perturbations reliably induce severe
misclassifications. This survey provides a comprehensive synthesis of adversarial attacks and defense mechanisms in modern
AI security. It introduces a structured taxonomy categorizing attacks into evasion, poisoning, and model inversion strategies,
evaluated across varying levels of attacker knowledge. Correspondingly, current defense techniques—including adversarial
training, anomaly detection, and gradient masking—are critically reviewed for their resilience against adaptive, real-world
adversaries. The survey further examines robustness benchmarking and success rate analysis frameworks, emphasizing the
gap between theoretical guarantees and practical deployment. By consolidating recent advances and persistent limitations,
this work identifies open research challenges and outlines emerging directions toward provably secure and trustworthy AI
systems for real-world applications.
This paper advocates for a forward-thinking approach that balances technical sophistication with human-centric principles, ensuring that adversarial deep learning evolves into a discipline not just of technical defense, but also of trust, transparency, and accountability.
Maisam Abbas, Ran-Zan Wang· IEEE Open Journal of the Com...· 0 citations
A detailed overview of the security risks associated with adversarial attacks is offered, including evasion attacks carried out at inference time, data poisoning that corrupts the training process, backdoor insertion that hides dormant triggers inside a model, and model inversion that leaks private information back out of a trained system.
Harsh Verma· International Journal of Sci...· 0 citations
The review explores the key adversarial attack classes: poisoning, evasion, model extraction, model extraction, model inversion, and membership inference and also white-box, black-box, and grey-box threat models.
Ujjwal Deshmukh· International Journal of Inn...· 0 citations
Many of the critical networks are now vulnerable to complex security threats, especially those launched by adversaries against the machine learning-driven security systems used by these networks. Such attacks take advantage of weaknesses in AI systems by perturbing the model with carefully designed perturbations, which result in misclassification of malicious content as benign, compromising the system's confidentiality, integrity, and availability. The adversarial threat is unlike traditional cyberattacks; it is dynamic, adaptive and can circumvent traditional intrusion detection capabilities. This paper provides an extensive literature review on the adversarial attack methods, detection and defence techniques of critical network infrastructures. This review includes peer-reviewed publications published between 2019 and 2024 from the leading academic databases such as IEEE Xplore, SpringerLink, ScienceDirect and Google Scholar. The total number of studies analyzed were 48, covering contributions in the fields of creating adversarial attack methods, machine learning and deep learning based detection methods, and mitigation techniques. The results indicate that adversarial attacks can be divided into the following categories: evasion attacks, poisoning attacks, and exploratory attacks, where some of the more sophisticated methods, including those based on gradient, optimization, and reinforcement learning, are very effective in evading security systems. Current solutions, however, suffer from limited real-time adaptability, cross-domain generalization ability, explainability and integration across the attack lifecycle. While there are several defence mechanisms proposed, such as adversarial training, anomaly detection, and input transformation, existing defences have difficulties in being adaptable in real time, cross-domain generalizable, explainable and suitable for certain phases of the attack lifecycle. The study highlights a number of critical research challenges such as the lack of a common defence framework, inadequate real-time detection capabilities, absence of a standardized data sets and poor ability to withstand adaptive adversaries. The paper suggests the creation of multi-strategic, adaptive, and real-time adversarial threat management systems that can sustain themselves in a heterogeneous network environment.
F. Okoye, Aghaizu Herman Chijioke, Shamsudeen Mohammed S.B· International journal of re...· 0 citations
A detailed empirical assessment of targeted adversarial vulnerability and defensive behaviour in a multi-class NIDS setting is presented and the results highlight long-standing, class-specific, robustness gaps and provide insights that could be used to design more robust intrusion detection systems.
A systematic analysis of 207 studies selected from 4447 records following the PRISMA 2020 guidelines, covering work published between 2020 and 2026 across cybersecurity and computer vision finds systems that are robust against adaptive adversaries, interpretable under operational constraints, and auditable in environments where AI accountability is a legal requirement.