Skip to content

Clinical Manuscript: Feasibility and Proof-of-Concept of the Rapid Dx Analyzer Structured Prompting Protocol: Achieving High Diagnostic Concordance in a Retrospective Case Series    

2026 · Journal of Medical Research and Reviews · Vol 5, pp. 116 · 0 citations

TL;DR

These preliminary findings suggest that structured prompt engineering may meaningfully improve LLM-assisted diagnostic reasoning and warrant further investigation through prospective, multi-center trials with blinded adjudication.

Abstract

Aim/Background: Diagnostic uncertainty persists as a major driver of preventable patient harm in clinical practice. This feasibility study examined whether a structured data-submission protocol — the Rapid Dx Analyzer and Clinical Decision Tool (R-DA) — could improve the reliability of a commercial large language model (LLM) in generating differential diagnoses for complex clinical presentations. Methods: A single-user, retrospective case series was conducted. Twenty non-consecutive clinical cases were selected from an institutional database using pre-defined complexity criteria (involvement of two or more organ systems, three or more differential diagnoses, ambiguous or conflicting data, or time-todiagnosis exceeding 48 hours). For each case, a structured prompt was submitted to Gemini 3.0 Pro via its web interface. The primary outcome was diagnostic concordance — defined as the model\'s top-ranked output matching the confirmed clinical diagnosis (established by biopsy, surgical findings, or definitive clinical course). The clinician providing input was not blinded to the final diagnosis. Results: Concordance between the R-DA-generated output and the confirmed diagnosis was observed in all 20 cases (100%; 95% CI [Clopper-Pearson]: 83.2%–100%). This result should be interpreted with caution given the small sample size and the lack of independent adjudication. Conclusion: These preliminary findings suggest that structured prompt engineering may meaningfully improve LLM-assisted diagnostic reasoning. The R-DA protocol warrants further investigation through prospective, multi-center trials with blinded adjudication. This study does not support claims of specialistlevel performance, but provides a hypothesis-generating foundation for future validation work

View source

Similar papers

Review Aug 2026

Clinical evaluation of a vision-language model for optimizing triage and clinical workflows in critical care.

OBJECTIVE Clinical deterioration in hospitalized patients is often preceded by subtle, dynamic physiological changes that are difficult to detect using intermittently charted electronic health record (EHR) data. Our objective was to evaluate the reliability, interpretability, and clinical relevance of a Vision‑Language Model (VLM)-based triage framework that analyzes physiological trend images, by comparing VLM-generated outputs with attending physician assessments as the expert clinical comparator. MATERIALS AND METHODS We conducted a single-center expert agreement pilot study including 100 adult patients with two hours of dynamic monitoring data across four vital signs (SpO2, RR, HR, BP). A structured prompt was developed using the Gemini 2.5 Flash model. Two independent reviewers assessed VLM outputs for clinical interpretation, artifact detection, and triage classification. We evaluated reviewer agreement using percent agreement and weighted Cohen's κ. A secondary risk-oriented analysis measured classification concordance and over-triaged cases as lower-risk classifications, while cases underestimating patient acuity were designated as higher-risk misclassifications. RESULTS The VLM demonstrated moderate to substantial agreement in triage classification with the attending physician and a low rate of under-triage. The VLM assigned the same patient acuity category as attending physician in 75 % of cases and underestimated acuity in 7 % of cases, compared with 14 % underestimation by the physician in training. DISCUSSION VLMs extend generative artificial intelligence capabilities by enabling image‑grounded clinical reasoning and offer signal‑processing capabilities for interpreting time‑stamped physiological trends. CONCLUSION The VLM demonstrated reliable clinical interpretation and an acceptable safety profile, however its integration into clinical workflows for early recognition of physiological deterioration and patient acuity assessment requires further rigorous evaluation and comparison to currently used track-and-trigger systems and patient monitoring methods.

I. Strechen, P. Krishnan, O. Kilickaya et al. · 0 citations
Review Open access Aug 2026

Informing the Development of a Conceptual Framework for a Diagnostic Radiology-Specific Patient-Reported Outcome Measure: A Systematic Review of Concepts, Instruments, and Gaps.

RATIONALE AND OBJECTIVES Radiology is central to clinical care, yet its value is described predominantly through metrics such as diagnostic accuracy, safety, and utilization, which do not capture the impact of imaging on patients. Despite growing evidence that imaging produces meaningful informational, emotional, physical, and logistical effects, radiology lacks a broadly applicable, psychometrically sound patient-reported outcome measure (PROM) for routine adult diagnostic imaging. The aim of this article is to synthesize the literature on patient-reported outcomes (PROs) and PROMs in adult diagnostic radiology, characterize current measurement approaches, and derive an evidence-based conceptual framework to inform development of CLARITY (Clinical and Life-impact Assessment of RadiologY), a new modular radiology-specific PROM. MATERIALS AND METHODS A prospectively registered, PRISMA-compliant systematic review searched MEDLINE, Embase, PsycINFO, and Web of Science from inception to February 2026. Eligible studies addressed adult diagnostic radiology and reported PROs relevant to imaging experience, uncertainty, result communication, or imaging-related burden. Structured narrative synthesis was performed, given methodological heterogeneity. RESULTS Of 21,693 records identified, 32 studies met eligibility criteria. Five recurrent patient-centered outcome domains emerged: informational outcomes, emotional outcomes, physical and procedural effects, burden, and perceived contribution to the healthcare trajectory. No existing instrument demonstrated content validity for routine adult diagnostic radiology; current approaches relied predominantly on generic quality-of-life tools, ad-hoc single-domain measures, and pathway-specific surveys. CONCLUSION The radiology patient-centered outcome space is sufficiently well-characterized to justify dedicated PROM development, yet no broadly applicable instrument has demonstrated content validity for routine adult diagnostic radiology. These findings provide the empirical and conceptual foundation for the development and future validation of the proposed CLARITY PROM program.

Rakhshan Kamran, Danial Aminaei, Benjamin Rehany et al. · 0 citations
Open access Jul 2026

Diagnostic capability of large language models in critically ill patients: a prospective single-centre study comparing ChatGPT, Claude, and Gemini with emergency physicians.

BACKGROUND Clinical decision-making requires integrating history, physical examination, laboratory, and imaging data. In the emergency department (ED), workload, time pressure, and cognitive burden may impair this process and affect decision quality. This study compares the diagnostic outputs of ChatGPT, Claude, and Gemini with those of emergency physicians in real-world ED cases. METHODS This prospective, single-centre observational diagnostic agreement study compared the stage-wise outputs of four Large Language Models (LLMs) (ChatGPT-4o, ChatGPT-5, Claude Opus 4.1, and Gemini 2.5 Pro) with those of emergency physicians in critically ill ED patients. Between 10 August and 10 September 2025, de-identified clinical data were entered into the models via their official web interfaces using standardised prompts. In the first stage, physicians and LLMs each generated five preliminary diagnoses based on vital signs and medical history. In the second stage, following physical examination and laboratory and imaging results, both refined their lists into three differential diagnoses. In the third stage, the physicians' final diagnosis was accepted as the reference, and each LLM was prompted to provide a final diagnosis. LLM preliminary and differential diagnoses were compared with those of the physicians at the corresponding stage, and LLM final diagnoses with the reference; the inclusion of the final diagnosis within earlier lists was also evaluated. Agreement was quantified using Cohen's κ; analyses were performed in R. RESULTS Of 389 screened patients, 180 were included (56.1% male; mean age 67 ± 15.9 years). Physicians contained the reference diagnosis within their top-5 preliminary and top-3 differential lists in 83.9% and 98.3% of cases, respectively, significantly exceeding every LLM (all p < 0.001). Final-diagnosis match rates were 67.2% [60.3-73.5] for ChatGPT-4o, 65.6% [58.7-71.9] for ChatGPT-5, 63.3% [56.3-69.9] for Claude Opus 4.1, and 59.4% [52.3-66.1] for Gemini 2.5 Pro (p = 0.16). Cohen's κ ranged from 0.575 (Gemini 2.5 Pro) to 0.656 (ChatGPT-4o), indicating moderate-to-substantial agreement, with no pairwise difference reaching significance. CONCLUSIONS The LLMs achieved moderate agreement with ED reference diagnoses in critically ill patients but were consistently outperformed by physicians at the early diagnostic phases. Despite final-diagnosis match rates of 59%-67%, their current diagnostic role in the ED remains limited.

İbrahim Günaydın, M. Yılmaz, Sinan Akpunar et al. · 0 citations
Open access Jul 2026

AI as a clinical decision support system: accuracy of an open GPT-model in emergency practice

Up to 30% of imaging examinations are deemed inappropriate, leading to unnecessary radiation exposure, emergency department (ED) overcrowding, and rising healthcare costs. While clinical decision support systems (CDSSs) such as ESR iGuide aim to improve imaging appropriateness, their clinical impact and user adoption remain limited. Large language models (LLMs) may offer a faster, more accessible alternative for radiologic decision-making. To evaluate the performance of OpenAccessGPT in reducing inappropriate radiological examinations requested from an ED by comparing its recommendations with ESR iGuide and assessing appropriateness and time efficiency. This retrospective, single-center, semi-quantitative cohort study was conducted at a tertiary hospital in Switzerland. A total of 201 consecutive ED radiology requests (March 2023–April 2024) met inclusion criteria (50 CT, 50 US, 50 X-rays, 51 MRI). Four hypotheses were tested: (1) OpenAccessGPT’s ability to reduce imaging requests; (2) concordance with ESR iGuide; (3) time efficiency; and (4) the rate of inappropriate examinations in clinical practice when confronted to ESR iGuide. Statistical analyses included proportion tests and paired t-tests ( p < .05). Mean patient age was 55.3 years (SD = 20.9); 45% were female. Eleven cases were excluded due to unavailable ESR iGuide scenarios. OpenAccessGPT advised against imaging in 6% of cases (12/201), with 33.3% agreement with ESR iGuide (5/12). Overall concordance was 78.9% (150/190), below the 90% threshold (χ² = 25.79, p < .001). GPT was significantly faster (21.6 s vs. 101 s; Δ = 79.6 s; t[189] = 12.8; p < .001). Institutional practice aligned with ESR iGuide in 82.6% of cases (χ² = 11.46, p < .001). OpenAccessGPT showed only limited ability to reduce inappropriate imaging. Despite substantial faster response times, its clinical reliability remained insufficient for standalone use.

Jérémie Arthaud, H. Thoeny, V. Grek et al. · 0 citations
Open access Jul 2026

Systemic bias in medical consensus: how flawed panels and lack of evidence can distort definitions and guidelines: the case of Vogt-Koyanagi-Harada disease : Perspective.

BACKGROUND The creation of diagnostic criteria for a disease is always challenging in medicine. Consensus meetings and expert's opinions are sometimes responsible for establishing useful and precise diagnostic criteria but, unfortunately, not all the consensus meetings are based on evidence-based conclusions. METHODS A perspective analysis of the historical evolution of diagnostic criteria for Vogt-Koyanagi-Harada. RESULTS We examine the historical evolution of diagnostic criteria for Vogt-Koyanagi-Harada disease, an autoimmune condition targeting melanocyte-containing tissues, primarily affecting the choroid and potentially involving the skin, ears, and meninges. Early diagnostic criteria proposed in the late twentieth century were considered inadequate, leading to an international consensus conference in 1999 and the publication of revised diagnostic criteria in 2001. However, these criteria contained major conceptual flaws, notably the combination of acute and chronic clinical features that rarely occur simultaneously. In addition, important diagnostic tools such as indocyanine green angiography were largely excluded due to prevailing local practices and skepticism regarding their use. These limitations hindered accurate diagnosis and delayed recognition of VKH as a disease with distinct acute-onset and chronic forms requiring different diagnostic frameworks and management strategies. Subsequent studies corrected some deficiencies but often generated complex criteria difficult to apply in routine practice. CONCLUSION We would like to emphasize the risks of poorly structured consensus processes and propose safeguards to ensure scientifically rigorous and clinically useful recommendations.

C. Herbort, Ioannis Papasavvas, A. A. El-Asrar et al. · 0 citations
Open access Jul 2026

Bedside Triage by Large Language Models in Acute Pancreatitis: A Scenario-Based Comparative Evaluation of GPT-4, GPT-5, and Gemini.

BACKGROUND Early decision-making in acute pancreatitis (AP) involves diagnostic confirmation, early severity triage, escalation thresholds, and initiation of guideline-concordant management under time pressure and incomplete information. Large language models (LLMs) may support structured bedside reasoning, but their clinical usefulness cannot be inferred from guideline knowledge alone. METHODS A cross-sectional, scenario-based comparative evaluation was conducted in January 2026 using 20 AP scenarios: 15 refined hypothetical vignettes and 5 de-identified, privacy-modified real-life case patterns. GPT-4, GPT-5, and Gemini received identical single-turn prompts. Model access was through OpenAI API gpt-4-0613, OpenAI API gpt-5, and Google Vertex AI Gemini 1.0 Pro; temperature was set to 0.0, and each prompt was repeated three times per model. Outputs were scored by two independent clinician-raters using a prespecified 1-5 ordinal rubric across guideline concordance, safety, actionability, and data-synthesis quality. Two senior board-certified surgeons independently generated expert reference pathways for comparison. RESULTS GPT-5 achieved the highest guideline concordance (4.28 ± 0.38) and safety (4.20 ± 0.45) profiles. GPT-4 provided the clearest stepwise actionability (4.15 ± 0.48), whereas Gemini showed the strongest data-synthesis quality (4.22 ± 0.52). With deterministic settings, internal consistency across three repeated runs was 100%. All models demonstrated clinically relevant failure modes, particularly unwarranted certainty under missing data; this occurred in 12/20 GPT-4, 7/20 GPT-5, and 15/20 Gemini outputs. CONCLUSION No model should be used as a stand-alone bedside decision-maker for AP. In this scenario-based early evaluation, GPT-5 was the most safety-aligned model, GPT-4 was the most operationally actionable, and Gemini was strongest for synthesis, but all require clinician oversight, prospective validation, and governance before clinical deployment.

Y. K. Çalışkan, Fatih Başak, Olgun Erdem · 0 citations