Large Language Models (LLMs) are now widely used to draft, revise, paraphrase, and polish text, making the detection of AI-generated writing increasingly difficult. This systematic literature review synthesizes peer-reviewed and high-quality studies published between 2023 and 2026 on AI-obfuscated, AI-refined, and humanized text. From 1,002 records, 26 primary studies were retained after screening and quality assessment. The review organizes the literature through a seven-dimensional taxonomy. Overall, the evidence shows that many detectors perform well on clean or in-distribution AI-text but become less reliable when the text is paraphrased, humanized, or collaboratively edited. The review also highlights recurring fairness concerns, especially for non-native English writers, and finds that current benchmarks often do not fully capture realistic mixed-authorship and adversarial settings. These results suggest that AI-text detection should be treated as one supportive signal rather than a stand-alone judgment, particularly in high-stakes academic or professional contexts.
Batyr Sharimbayev, S. Kadyrov· Engineering, Technology &...· 0 citations
Test set results show that Decoding-Enhanced Bert with Disentangled Attention (DeBERTa) achieves the highest macro F1 − Score of 85.48%, surpassing the previously top-ranked Multi-Task Learning (MTL) system, which attains a macro F1 of 83.07%.
Batyr Sharimbayev, S. Kadyrov· Journal of Advances in Infor...· 0 citations