Towards a Code-Based Assessment of Computational Thinking Skills
Abstract
Computational thinking (CT) skills are central to higher-education programming, yet assessing them at scale is challenging. Automated source-code analysis is a promising path, but its validity at university level is still poorly understood. We examined how far code-based metrics can capture traces of two CT skills: abstraction and algorithmic thinking. In a controlled study, 44 students enrolled in a Data Structures course completed four programming exercises designed to elicit evidence of the two target computational thinking skills. Each submission was assessed using three independent sources: an automated score based on structural code metrics, a rubric-based score by experts, and a CT self-efficacy scale. The three sources were compared using Spearman correlations, quartile analyses, and qualitative inspection of fourteen concordant and divergent cases. At the per-student level, the three sources did not converge for either skill. Per-exercise, one converged with expert evaluation, while another showed that the sources captured distinct, complementary facets of the skill—an exercise-design effect. Qualitative cases reinforced this view, illustrating what each source reads: from structural shape to semantic correctness. Automated code-based assessment of CT can provide informative structural evidence, but works best as one element of a multi-source approach in which structural signals complement, rather than replace, expert evaluation.