AI-Powered Automated Speaking Scoring and Human Evaluation: A Comparative Study of EFL Learners’ Oral Proficiency
Recent advances in artificial intelligence have led to the development of automated speaking assessment systems capable of evaluating oral proficiency with increasing accuracy. This study compares Artificial Intelligence (AI) powered automated speaking scoring and human evaluation in assessing the oral proficiency of 74 Saudi English as Foreign Language (EFL) learners. Participants completed a picture-description speaking task, which was evaluated by both Claude and two trained human raters using identical holistic and analytic speaking rubrics. The study examined holistic speaking scores and five analytic dimensions: coherence, cohesion, content development, grammar, and vocabulary. Statistical analyses included descriptive statistics, comparative analyses, Intraclass Correlation Coefficients (ICC), Weighted Kappa coefficients, and Bland-Altman analysis. Results showed no significant difference between AI-generated and human-assigned holistic speaking scores. Similarly, cohesion and grammar demonstrated strong similarity between the two assessment approaches. However, significant differences were observed for content development, vocabulary, and coherence, with Claude consistently assigning slightly higher scores than the human raters. Agreement analyses revealed good-to-strong agreement across both holistic and analytic assessments. The findings suggest that AI-powered speaking assessment can produce evaluations broadly comparable to human judgments while demonstrating strong consistency across multiple dimensions of oral proficiency. The study supports the potential of AI-assisted speaking assessment as a reliable complement to human evaluation in EFL contexts.