Large language models and natural language processing applications in radiology: a systematic review of Q1 journal studies
Abstract
Aims: To systematically review large language model (LLM) and natural language processing (NLP) studies published in first-quartile (Q1) clinical radiology journals, focusing on methodological quality, model implementation, and comparative performance. Methods: A systematic search of PubMed and Scopus was conducted to identify original studies involving LLMs or transformer-based NLP systems published in Q1 clinical radiology journals through June 20, 2025. Eligible studies were screened and assessed for methodological characteristics, including dataset type, involving imaging modality (if any), model used, model accessibility, prompt disclosure, and handling of stochasticity. Human-LLM/NLP and LLM/NLP-LLM/NLP performance comparisons were extracted. Results: Fifty-six studies were included, most published in 2024-2025. Proprietary models such as GPT-4 and GPT-4o were most frequently evaluated. Real-world clinical data were used in 62.5% of studies, but only 10.7% reported a power analysis, and 39.1% addressed stochasticity. Prompt engineering was reported in 41.9% of studies. In 455 human-LLM/NLP comparisons, LLMs/NLPs outperformed humans in 54 cases, while humans outperformed in 79; most results (70.8%) were ties. Among 3,164 valid LLM/NLP-LLM/NLP comparisons, GPT-4o had better performance than earlier models. Conclusion: LLMs/NLPs demonstrated performance comparable to radiologists in many text-based tasks but remain inconsistently evaluated. Methodological limitations, including lack of power analysis, incomplete reporting, and under-addressed stochasticity, hinder robust assessment. Greater transparency, standardized evaluation protocols, and inclusion of diverse clinical settings are essential for reliable integration into radiology practice.