Skip to content
Review

LLMs for Survey Text Analysis - A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis

Aug 2026 · 0 citations
Computer Science

TL;DR

Overall, the findings suggest that LLMs can approximate human coding in this case-specific setting, particularly at the coding level, and may serve as a scalable support tool for inductive qualitative analysis.

Abstract

Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in inductive content analysis remains limited. This study compares human and LLM-based inductive coding of open-ended survey responses from 903 answers across six variables from a European PhD student survey. Five human coders performed inductive content analysis following a standardized coding scheme, while an LLM (GPT-5.4) conducted the same task using an established prompting procedure. Agreement between human and LLM outputs was assessed using the Adjusted Rand Index (ARI). Results showed an alignment between humans and the LLM, with ARI values of 0.61 for coding and 0.54 for theme generation. These values were close to the internal consistency of coding and theme results within humans (ARI = 0.68) and the LLM (ARI = 0.76). Agreement varied widely across variables, with low within-entity consistency consistently linked to low between-entity agreement, underscoring the role of data characteristics and individual performance in reliability. Overall, the findings suggest that LLMs can approximate human coding in this case-specific setting, particularly at the coding level, and may serve as a scalable support tool for inductive qualitative analysis.

View source

Similar papers

Open access Aug 2026

Evaluation of Inductive Coding with LLMs

Large Language Models (LLMs) like ChatGPT are reshaping qualitative research by offering data-driven analysis. Recent studies focus on labeling and classifying qualitative data, yet the generative process behind code system development is underexplored. This study investigates how ChatGPT constructs coding systems from interview data under iterative code engineering. The resulting code system is systematically compared to an inductive coding framework derived by content analysis—quantitatively by SBERT analysis, Jaccard index, and network analysis, and qualitatively. The results of NLP-based analysis show that there are mostly high cosine similarities in the one-to-one mapping between ChatGPT codes and content analysis. The overall Jaccard index is low, indicating limited overlap between the coding units and, consequently, that the two coding systems often did not refer to the same content. However, the mapped codes showed substantial variation: while nearly half exhibited high overlap, others showed no overlap. Network analysis shows that network density is comparable across the two systems, indicating a similar level of interconnectedness among codes despite differences in network size. Finally, iterative prompt engineering could create a stable code system, but ChatGPT lacks in domain-specific coding addressing the research questions adequately, warranting further investigation across larger datasets.

Leoni Dörfel, Rieke Ammoneit · 0 citations
Open access Aug 2026

Don’t Believe the Hype: Methodological Approaches for Applying LLM-Assisted Content Analysis to Reported Speech in Journalism

To better understand how journalists represent sources, it is necessary to systematically study the use of reported speech in news coverage. This paper presents a method for LLM-assisted content analysis to identify and classify reported speech in Dutch newspapers automatically. The study evaluates a three-step procedure utilising role-based instructions to prompt the model as a professional journalist. First, a codebook for identifying citation structures and source types was developed with LLM support and then manually verified. Second, inter-coder reliability between human coders and the LLM was assessed on a representative sample of Dutch news articles using a human-in-the-loop validation approach. Third, the prompt-engineered LLM was used to code a large corpus spanning seven decades (1950–2024). Manual verification of 16,689 citations shows a weighted F1-score of 0.75, which aligns with recent benchmarks for high-capacity models performing complex journalistic coding. While human oversight remains the benchmark for reliability, due to issues such as repeated citations that were given as examples in the used prompts and representational bias, LLM-based systems perform sufficiently well for large-scale analyses of journalistic source use. The paper concludes that hybrid human–AI workflows provide a practical bridge between traditional rule-based approaches and new generative models, offering scalable and cost-effective methods for studying source representation in journalism.

J. de Cooker · 0 citations

Have Large Language Models Improved Research Methodology?

Whether contemporary LLMs can reproduce the research outcomes of a fully documented human study: a 1991 article that identified dermatophytosis (ringworm) in historical fine art was evaluated.

Fredric Narcross, Robert Marks · 0 citations
Review Open access Aug 2026

Large Language Models in Peer Review: Decision Alignment, Review-Text Characteristics, and Human–AI Aggregation at ICLR 2025

Large language models (LLMs) are increasingly employed in scholarly peer review, yet their suitability as autonomous evaluators remains uncertain. Using the ICLR 2025 review process, this study compares 2401 human reviews with 7203 reviews produced in separate, context-isolated API runs using Claude Sonnet 4.5, GPT-5.2 Thinking, and Gemini 3 Pro Preview across decision agreement, review-text characteristics, inter-model consistency, and human–AI aggregation. Raw LLM scores showed systematic leniency and score compression. A 0.1-point grid search identified thresholds of 6.2, 6.3, and 6.7 for Claude, GPT, and Gemini, respectively; repeated stratified cross-validation reproduced these thresholds. When applied without retuning to a stratified balanced sample of 300 ICLR 2024 papers, decision-agreement accuracy was 0.927, 0.913, and 0.930. Independent human coding of research type and primary field showed substantial pre-adjudication agreement (Cohen’s kappa = 0.774 and 0.714), and the recalculated analyses did not support H3. Review-text indicators showed similar structural completeness across sources but uneven critical-section length; these descriptive measures do not establish review quality. Human-containing aggregation rules showed higher agreement with conference decisions than corresponding AI-only rules, without establishing independent review quality or causal complementarity. A textual-overlap check found very low exact eight-gram containment, and manual inspection of the highest-similarity 1% found shared manuscript content or domain terminology rather than reviewer-specific evaluative language; possible prior exposure nevertheless could not be excluded.

Zhihe Yang, Xiao-Yue Zhou, Hong-Sa Wang et al. · 0 citations
Open access Aug 2026

EduFairBench: reproducible evaluation of large language models for educational assessment

Large language models (LLMs) are increasingly used to evaluate open-ended educational responses. However, their performance is often assessed using aggregate metrics that provide limited insight into prediction stability, uncertainty, error patterns, and feedback quality. This study presents EduFairBench, a reproducible evaluation protocol designed to characterize LLM behavior across short-answer assessment and automated essay scoring using open educational benchmarks. The protocol combines repeated inference, majority-vote consolidation, uncertainty estimation, error analysis, and structural evaluation of generated feedback within a unified experimental framework. Experiments were conducted on SciEntsBank, Beetle, and ASAP2, comprising 2,000 student responses and 10,000 independent LLM inferences. The results showed moderate predictive agreement with human assessment while revealing substantial differences between nominal and ordinal evaluation tasks. Repeated inference demonstrated high internal stability across benchmarks, although systematic errors remained in semantically adjacent categories, indicating that prediction consistency does not necessarily imply correctness. Feedback quality varied by task type, with longer textual contexts yielding more specific and pedagogically structured explanations. These findings demonstrate that evaluating educational LLMs requires complementary analyses beyond conventional performance metrics. EduFairBench provides a reproducible methodology for jointly analyzing predictive performance, robustness, uncertainty, and feedback quality, providing a comprehensive methodological framework for the rigorous evaluation of LLM-based educational assessment systems.

W. Villegas-Ch., Aracely Mera-Navarrete, Fernando Zúñiga-Tello et al. · 0 citations