Detecting AI-Generated Text: Mechanisms, Robustness, and the Limits of Reliable Detection
Abstract
Large language models (LLMs) have become capable of producing human-like prose, institutions ranging from universities to publishers have adopted automated AI-text detectors — most visibly Turnitin's AI writing indicator, GPTZero, and similar tools — as a control against misrepresenting machine-generated content as human work. This paper examines how these detectors work, the empirical and theoretical evidence on their reliability, and the documented techniques that allow AI-generated text to evade them. Drawing on the peer-reviewed and preprint literature, we describe three detector families (zero-shot statistical detectors, trained classifiers, and watermarking schemes), summarize adversarial results showing that paraphrasing attacks can collapse watermark detection true-positive rates from above 99% to under 10%, and review evidence that current detectors produce systematically higher false positive rates for non native English writers — in one widely cited study, 61.3% of TOEFL essays were misclassified as AI generated versus near zero misclassification of native-speaker essays. We conclude that AI-text detection, in its current form, cannot serve as a sole, dispositive basis for academic-integrity decisions, and we propose a set of institutional and technical recommendations — process-based evidence, disclosed AI-use policies, watermarking at the model level, and human-adjudicated review — that better address the underlying problem than detector accuracy alone can.