As computation becomes more central to physics education, creating scalable methods to assess authentic computational thinking (CT) in students remains a critical challenge. While student-written responses capture nuanced reasoning, they are difficult to evaluate at scale. In this study, we investigated the use of Large Language Models (LLMs) to analyze students'written explanations of computational physics problems on a pre- and post- semester survey. By first establishing a human-coded baseline, grounded in CT literature, we identified significant growth in Data Practices and Computational Problem-Solving Practices. When given the same responses, an LLM successfully mirrored the human evaluations and scaled up the detection of these key trends across a large dataset. Notably, both human raters and the LLM struggled to reliably evaluate more complex constructs such as Systems Thinking. Overall, this study demonstrates that LLMs offer a viable method to scale the evaluation of students'CT in large-enrollment physics courses
S. Savage, Anand Shanker, Grace Michlitsch et al.· 0 citations
This study examines Artificial Intelligence (AI)-generated physics solutions from two connected perspectives: how prompt design shapes these solutions and how students can be prepared to critique them. Using a rotational-mechanics problem, we adapted a problem-classification framework to examine prompt variations, evaluating OpenAI's o4-mini responses with the Minnesota Assessment of Problem Solving (MAPS) rubric. Well-specified prompts improved solution completeness; underspecified and multimodal prompts exposed weaknesses in physics reasoning and correctness. In the student-evaluation phase, 24 introductory physics lab groups evaluated an o4-mini solution to this problem after either independently solving a related problem or critiquing its AI-generated solution with MAPS-based reflection questions. Problem-solving-only groups exhibited uncritical or misconception-based critiques; MAPS-guided groups identified more expert-aligned issues, including skipped numerical procedures and undefined notation. Together, our findings contribute to physics education research by showing how AI-generated solutions can ground both model-reasoning benchmarks and improved student critique of that reasoning through MAPS-based reflection.
N. Borse, Amir Bralin, S. Savage et al.· 0 citations