Skip to content

Author

Huilin Chen

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

COMPARATIVE EVALUATION OF THREE LLMS FOR CEFR-ALIGNED READING PASSAGE AND ITEM GENERATION

Research Objectives: This study compares three large language models (Deepseek, GPT-4, Baidu Ernie 4.0) in automatically generating CEFR A1–C2 reading comprehension passages and test items. Methodology: Each model produced 72 passages (narrative, expository, argumentative, instructional) and 120 four-option multiple-choice items covering nine CEFR skill descriptors. Five certified language assessment experts independently rated all materials on nine quality dimensions using 5-point Likert scales. Inter-rater reliability was excellent (Fleiss’ κ = 0.82–0.91). Data were analysed with Kruskal-Wallis H tests and post-hoc Dunn–Bonferroni corrections. Findings: Passage and item quality were strong up to B2 level (M > 4.2/5.0 across models), but declined notably at C1–C2 (M = 3.08–3.78). Higher-order skills (e.g., rhetorical purpose evaluation, implicit attitude inference) scored significantly lower (p < .001). Deepseek outperformed GPT-4 on CEFR alignment (p = .002) yet remained 0.5–0.7 points below human-authored benchmarks. Research Outcomes: Current LLMs can reliably generate psychometrically sound reading materials up to B2 level, but remain inadequate for fully automated high-stakes C1–C2 assessment without intensive human post-editing. Future Scope: Enhance LLM capabilities for rhetorical complexity, cohesion manipulation, and higher-order item design targeting advanced proficiency levels.

Huilin Chen · 0 citations