Skip to content
Open access

Predicting Authorship Attribution in AI-Generated vs Human-Authored Texts: A Corpus-Based Study of Syntactic Complexity in Personal Statements

Aug 2026 · Dialogica · 0 citations · 40 references

Abstract

Generative AI complicates the use of personal statements as evidence of applicants’ individual voice. This corpus-based quantitative study examined whether syntactic-complexity measures distinguish human-authored from ChatGPT-generated personal statements for business and economics fellowship applications. The corpus comprised 50 publicly accessible human-authored texts and 50 texts generated by GPT-4o from a single prompt. Lu’s L2 Syntactic Complexity Analyzer yielded nine variables: word, sentence, and clause counts; dependent clauses per clause (DC/C) and T-unit (DC/T); complex T-units per T-unit (CT/T); coordinate phrases per clause (CP/C) and T-unit (CP/T); and T-units per sentence (T/S). Descriptive statistics, one-way MANOVA, follow-up ANOVAs, and discriminant function analysis were applied. The multivariate effect of text type was significant, Wilks’ Λ = .075, F(9, 90) = 123.51, p < .001, partial η² = .925; the discriminant function was also significant, χ²(9, N = 100) = 242.31, p < .001. Human texts were much longer (M = 915.80 vs. 150.68 words) and showed higher DC/C, DC/T, CT/T, and T/S values; CP/C and CP/T did not differ. Thus, text length and selected subordination and T-unit measures distinguished the corpora. Because the texts were not length-matched and the design used one model, one prompt, and no cross-validated classification, the findings are corpus-specific and cannot support a general-purpose AI-text detector.

Read PDF