Evaluating Command-Line Interface-Based Agentic Large Language Model Coding Tools for Non-English Thematic Analysis: A 3 x 3 x 3 Factorial Study of Models, Prompts and Stochastic Variability
Abstract
Command-line interface (CLI)-based agentic coding tools enable large language models (LLMs) to autonomously read local data files, write and execute analysis scripts, and generate outputs, but whether they can reliably perform thematic analysis of non-English educational data remains understudied. This 3 × 3 × 3 factorial experiment compared three frontier LLMs (Claude Opus 4.6, GPT-5.3-Codex, and Gemini 3 Pro Preview) across three prompt tiers (detailed, structured, and exploratory) with three repetitions per condition, yielding 27 runs. Each run processed 89 anonymized student evaluations (Likert ratings and open-ended Thai comments) from a Year 3 preventive dentistry course over three academic years. Outputs were evaluated against 10 required items covering statistical computation, thematic analysis, and visualization, with the generated scripts inspected to classify categorization as direct reading or keyword matching. Statistical computation and overview charts each passed 93% (25/27), and detail charts passed in all 27 runs, with failures confined to exploratory prompts. Thematic reading was the primary difference. Opus and Gemini read Thai directly while GPT usually generated keyword-matching scripts. Structured prompts achieved the highest thematic reading pass rate at 78% (7/9), followed by detailed at 67% (6/9) and exploratory at 33% (3/9). Identical configurations produced different outputs across repetitions, with theme counts ranging from 4 to 31. CLI-based agentic coding tools can perform thematic analyses of Thai-language educational feedback, but natural language processing output reliability depends on model selection, prompt specificity, and stochastic variability. Multiple independent runs are recommended to assess consistency.