This paper systematically synthesizes 27 studies of attacks against MGTD and the available evidence on corresponding defenses, categorizing existing research into four major types of evasion strategies: watermark attacks, paraphrasing attacks, prompt-based attacks, and adversarial-text attacks, and summarizes the available defense evidence.
Abstract
With the explosive growth of large language models (LLMs), research on machine-generated text detection (MGTD) has also proliferated. Alongside these developments, a wide range of attack algorithms targeting MGTD systems have emerged. While previous studies have surveyed detection techniques, few have examined the dynamic interplay between attack and defense. Following PRISMA 2020, this paper systematically synthesizes 27 studies of attacks against MGTD and the available evidence on corresponding defenses. We categorize existing research into four major types of evasion strategies: watermark attacks, paraphrasing attacks, prompt-based attacks, and adversarial-text attacks, and summarize the available defense evidence. Furthermore, to better understand the practical implications of these methods, we compile the reported performance results of attack and defense techniques across different detectors. Finally, we highlight the current challenges in this area and outline potential future research directions. A companion repository containing the categorized literature, paper links, and available code, data, and project repositories is provided at https://github.com/AIGC1999/A-Survey-of-Anti-Detection-Techniques-for-Machine-Generated-Texts.
Problem-space evasion attacks have exposed critical weaknesses in machine learning-based malware detectors; yet, their evaluation remains fragmented across models, datasets, and attack methodologies, often neglecting domain-specific requirements such as executability and functionality preservation. We address this gap...
Mashal Zainab, Salijona Dyrmishi, Hamid Bostani et al.· 0 citations
Machine learning-based network intrusion detection systems (ML-based NIDS) are vulnerable to adversarial evasion, where malicious samples are perturbed to evade detection and be misclassified as benign. Despite growing research on adversarial attacks and defenses for ML-based NIDS, comparative evaluations of multiple a...
Two universal tool-based defenses are introduced: Attacker Tool Filtering, which uses anomaly detection to identify and remove suspicious tools, and Normal Tool Recalling, a white-box method that restores the agent's original toolset prior to planning.
It is demonstrated that clean-text performance is not a reliable predictor of adversarial robustness, and the results underscore the necessity for architecture-specific defences and frame smishing detection as an adversarial cybersecurity challenge rather than a static classification task.
Denzel Chiuseni, A. Bahizire, Silva Hama et al.· 0 citations
A benchmark for this vulnerability in LLM-based resume screening is introduced: 463 job-candidate pairs drawn from a 14-domain corpus, with the evaluated sample covering 13 domains, attacked through a taxonomy of four attack types and four injection positions.
Hong-Lin Mu, Jinghao Liu, Kai-Yang Wan et al.· International Journal of Mac...· 4 citations· ⚡1
The solution, LSABRE, is a multi-LLM framework that improves robustness across various attacks, maintaining 86% detection accuracy even under strong adaptive adversarial attacks, and proposes a robust multi-LLM defense architecture designed to preserve detection reliability under adaptive adversarial conditions.