GPT-4’s performance on the UKMLA improved significantly over 18 months, suggesting that iterative vendor-side model updates enhance clinical reasoning capabilities, and may serve as a supplementary educational tool for medical students and clinicians.
Abstract
Background Generative Artificial Intelligence (GAI) models, such as GPT-4, have been extensively studied for their integration into medical practice and education. GPT-4 has demonstrated excellent performance on medical licensing examinations, including the United Kingdom Medical Licensing Assessment (UKMLA). However, the field lacks longitudinal data on whether such performance is stable or varies over time. Theoretically, iterative improvements updated by the vendor should enhance performance, but empirical evidence of such longitudinal trends remains limited. Given that GPT-4 undergoes periodic vendor-side updates, we aimed to analyse the categorical, time-spaced performance of GPT-4 on the UKMLA to characterise how its performance changes over time in a medical context. Methods Two publicly available UKMLA papers were fed into GPT-4 at two different time points, June 2023 and December 2024. 191 questions were provided with and without multiple-choice options to assess GPT-4’s clinical competence. McNemar’s test was performed to evaluate changes in GPT-4’s performance over time, comparing domain-specific questions. Results GPT-4’s accuracy improved noticeably between the two rounds (MCQ: 88.0 to 93.7%, p = 0.027; non-MCQ: 68.1 to 81.7%, p = 1.00). Single-step accuracy rose from 73.1 to 82.3%, and multi-step from 57.4 to 80.3% without MCQ. GPT-4 showed improved accuracy from Round 1 to Round 2 for both single-step and multi-step questions, with MCQ-prompted responses consistently outperforming non-MCQ responses (up to 95.1% accuracy for multi-step MCQ questions in Round 2). GPT-4’s performance improved across all question categories from round one to round two, most notably in management questions without MCQ options (+23.30%), though these differences were not statistically significant. Discussion and conclusion GPT-4’s performance on the UKMLA improved significantly over 18 months, suggesting that iterative vendor-side model updates enhance clinical reasoning capabilities. These findings indicate that GPT-4 may serve as a supplementary educational tool for medical students and clinicians; however, the underlying drivers of performance changes remain opaque, and such tools should be deployed with structured oversight to prevent overreliance.
Newer models may be more robust, but multiple-choice accuracy remains an incomplete measure of clinical reasoning robustness, as all evaluated LLMs showed a significant robustness gap after NOTA replacement.
A. Soejima, F. Kitano, D. Ichikawa et al.· medRxiv· 0 citations
Clinical deterioration unfolds through coupled, partially observed trajectories, not a single diagnostic label. We introduce PGP-Clinical-TimeKAN, a trajectory-first framework for joint probabilistic forecasting of multivariate physiology. It combines missingness-aware temporal encoders, a soft organ-system prior, pati...
Weizhi Nie, Rihao Chang, Weijie Wang et al.· 0 citations
Medical time series (MedTS), including electrocardiograms (ECG), electroencephalograms (EEG), photoplethysmography (PPG), and vital-sign recordings, are central to clinical diagnosis and health monitoring. As large language models (LLMs) have advanced, a growing body of work has examined how their reasoning, generation...
Yu Han, Cigdem Beyan, Xiang Zhang et al.· 0 citations
Background: Large language models (LLMs) have demonstrated strong performance on standardized medical examinations, with recent studies reporting performance approaching or exceeding that of senior medical residents. However, examination accuracy alone does not establish how models arrive at their answers or the relati...
F. Gafoor, M. Syed, M. Halai et al.· medRxiv· 0 citations
Background Laboratory-based frailty indices (FI-Lab) have shown promise in geriatric medicine research. We aimed, in diverse samples, to determine the optimal construction of an FI-Lab for acute care and evaluate its validity as a measure of latent health status across the adult life span. Methods and findings Our retr...
H. L. Ellis, Peter Hanlon, Liam Dunnell et al.· PLoS Medicine· 0 citations
Abstract Background Following the transition of the USMLE Step 1 to pass/fail scoring, Step 2 Clinical Knowledge (CK) has become a primary metric in residency screening. This study evaluates multiple approaches to identify performance patterns that may support individualised longitudinal advising and targeted preparati...
B. Boateng, K. Clemmons, Lindsey Sward et al.· Medical Education Online· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.