Skip to content
Open access

Syntactic Complexity in AI-Generated vs. Human-Authored Linguistic and Literary Texts

Asia A. Alheety Meethaq Khamees Khalaf H. Mohammed
Jul 2026 · Arab World English Journal · Vol 17, pp. 109-125 · 0 citations · 13 references

TL;DR

The results indicate that the complexity of syntax is genre-based and not source-based and in general, the discipline genre had a more significant effect on syntax variation than the authorship source.

Abstract

This paper examines how the syntactic complexity of academic writing is affected mainly by the source of authorship (AI-generated or human) or by the genre of disciplinary writing (linguistic or literary). The primary purpose is to test the syntactic-complexity differences between these variables and to establish the degree of influence of genre conventions on structural variation. The importance of the research is that it adds to the existing discussions about AI as a phenomenon in academic writing and, specifically, whether AI-generated texts are capable of syntactically reproducing the specific norms of a specific discipline of writing. To this end, the comparative corpus-based design was used. The sample consisted of 20 introduction sections: equal numbers of linguistic and literary texts and equal numbers of human-authored AI-generated texts. The Second Language Syntactic Complexity Analyzer (L2SCA) extracts fourteen syntactic complexity measures, which include length of production unit, subordination, coordination and phrasal sophistication. The results indicate that the complexity of syntax is genre-based and not source-based. Although no differences were found to be constant in both AI-generated and human academic introductions in linguistic data, there were much higher levels of subordination in literary texts of human origin. In general, the discipline genre had a more significant effect on syntax variation than the authorship source. The paper suggests the implementation of genre-sensitive models to assess AI-written academic texts and recommends additional studies that would use a bigger sample and discourse analysis.

Read PDF

Similar papers

Open access 2026

Characterization and Mechanisms of Lexical Complexity in AI-Generated Texts: A Comparative Corpus-Based Study

: Based on a corpus-based methodology, this study analyzes the intrinsic reasons for the high level of lexical complexity observed in Artificial Intelligence Generated Content (AIGC). The research compares 24 English argumentative essays written by AI with 24 second-language (L2) learner essays reaching the IELTS Writing Task 2 Band 7 level. Under controlled conditions of identical genre and topic, the study performs quantitative statistics across three dimensions: lexical sophistication, semantic abstraction, and information density. Statistical results indicate that the frequency of advanced vocabulary in AI texts is significantly higher, approximately 2.3 times that of human texts. The proportion of abstract nouns reached 9.14%, far exceeding the 3.32% found in human texts, suggesting that AI expressions tend toward nominalization and conceptualization. Regarding overall information organization, the lexical density of AI texts was 69.9%, also surpassing the 60.3% of human texts, reflecting a stronger tendency for information condensation and phrasal structures. The analysis points out that the complexity of AI text primarily stems from its mechanism of selecting vocabulary based on probability distributions. This mechanism favors longer words, abstract nouns, and words with high semantic content, thereby forming a highly compact linguistic surface. Such complexity is essentially a formal feature at the statistical level and is not entirely equivalent to the proficiency levels corresponding to human L2 acquisition. These findings provide empirical references for AI text identification, the refinement of writing evaluation standards, and L2 writing pedagogy.

Lulu Chen · 0 citations
Review Open access Jul 2026

Linguistic Features of AI-Generated Academic Texts and the Role of Human Editing

The rapid spread of large language models (LLMs) has significantly transformed academic writing practices and actualized discussions about authorship, language quality, and academic integrity. At the same time, diachronic changes in academic discourse during the active implementation of generative artificial intelligence remain insufficiently studied. The study combines a systematic literature review with a corpus diachronic analysis of authentic academic annotations, allowing us to trace long-term trends in the development of academic discourse. The aim of the work is to identify linguistic changes in academic writing during 2015–2025 and determine the role of human editing in quality assurance of AI-assisted scientific texts. The research material was a corpus of 870 English-language annotations of scientific articles indexed in the Scopus database in the field of arts and humanities. Quantitative linguistic analysis was carried out using Python tools. Indicators of lexical density, lexical diversity, syntactic complexity, average sentence length and frequency of cohesive markers were analyzed. A statistically significant increase in lexical density and frequency of cohesive markers has been revealed, indicating an increase in information compression and explicit discursive organization of texts. Indicators of traditional lexical diversity and syntactic complexity remained relatively stable. The observed trends are consistent with characteristics described in AI-assisted writing studies; however, the study design does not allow them to be directly related to the use of large language models. Human editing remains a key factor in ensuring factual accuracy, discursive coherence, lexical enrichment, and academic integrity. The results can be used to develop practices for the responsible use of generative AI in academic communication.

T. Nedashkivska, I. Varvaruk, M. Podoliak et al. · 0 citations
Open access Jul 2026

Human vs AI-Generated Texts in Language Learning: A Linguistic Comparison

The findings show that AI-generated texts exhibit greater lexical diversity and syntactic complexity; however, they often exhibit structural uniformity, overuse of cohesive devices, and limited pragmatic depth, and should not replace professionally designed educational materials.

V. Smaglii, T. Korolova, Svitlana Yukhymets et al. · 0 citations
Open access Jul 2026

Texts Generated by Artificial Intelligence: Structure and Semantics

It was concluded that texts generated by artificial intelligence constitute a separate linguistic phenomenon with its own set of characteristics, which requires a special typology and a flexible, updatable analysis methodology.

L. Kravets, Viktória Stefuca, N. Libak et al. · 0 citations
Open access Aug 2026

Predicting Authorship Attribution in AI-Generated vs Human-Authored Texts: A Corpus-Based Study of Syntactic Complexity in Personal Statements

Generative AI complicates the use of personal statements as evidence of applicants’ individual voice. This corpus-based quantitative study examined whether syntactic-complexity measures distinguish human-authored from ChatGPT-generated personal statements for business and economics fellowship applications. The corpus comprised 50 publicly accessible human-authored texts and 50 texts generated by GPT-4o from a single prompt. Lu’s L2 Syntactic Complexity Analyzer yielded nine variables: word, sentence, and clause counts; dependent clauses per clause (DC/C) and T-unit (DC/T); complex T-units per T-unit (CT/T); coordinate phrases per clause (CP/C) and T-unit (CP/T); and T-units per sentence (T/S). Descriptive statistics, one-way MANOVA, follow-up ANOVAs, and discriminant function analysis were applied. The multivariate effect of text type was significant, Wilks’ Λ = .075, F(9, 90) = 123.51, p < .001, partial η² = .925; the discriminant function was also significant, χ²(9, N = 100) = 242.31, p < .001. Human texts were much longer (M = 915.80 vs. 150.68 words) and showed higher DC/C, DC/T, CT/T, and T/S values; CP/C and CP/T did not differ. Thus, text length and selected subordination and T-unit measures distinguished the corpora. Because the texts were not length-matched and the design used one model, one prompt, and no cross-validated classification, the findings are corpus-specific and cannot support a general-purpose AI-text detector.

M. Ayadi · 0 citations
Open access Aug 2026

Evaluating the Role of Artificial Intelligence in Supporting Usage - Based Approaches to Grammar

This study investigates grammatical trends in texts generated by artificial intelligence and human learners. The study puts to the test a fundamental principle of usage-based grammar: language is learned through repeated exposure to patterns. A direct comparison is conducted between AI-generated writings and language learners' essays. Quantitative approaches count words, sentences, and grammatical errors. Qualitative analysis detects trends in sentence structure and specific qualities such as past tense. Finding out if AI models adhere to usage-based grammar rules is the aim. Comparing the two groups' mistake types is another objective. The results show that whereas human writing varies, AI output is very constant. Almost no grammatical errors were found in AI articles, according to the study. Expected errors in human texts include omissions and overgeneralizations. The findings also demonstrate that AI makes greater use of components like the past tense and plurals. These studies demonstrate that the outcomes of usage-based learning are operationally replicated by AI. The results of training the model on massive amounts of data are consistent and precise. The ongoing process of language acquisition is reflected in human output. The study comes to the conclusion that AI is a powerful instrument for confirming frequency-based linguistic theory.but does not model the human cognitive journey. Future research should investigate different AI models and learner proficiency levels

Assis. lect. Batool Abdul-Mohsin Miri · 0 citations