Skip to content
Preprint

Improving O-RADS Risk Stratification from Ultrasound Reports: A Comparative Evaluation of Hybrid versus End-to-End LLM Reasoning Strategies

Aug 2026 · 0 citations · 30 references
Computer Science

TL;DR

The feature-based hybrid LLM architecture that separates clinical feature extraction from deterministic guideline execution enables highly accurate, reliable, and interpretable automated O-RADS classification, providing a promising approach for standardized, guideline-based clinical decision-making.

Abstract

Background: Automating clinical guideline-based decision-making with large language models (LLMs) remains challenging because of reliability, hallucination, and limited interpretability. We compared the performance of LLMs and reasoning strategies for automated Ovarian-Adnexal Reporting and Data System (O-RADS) classification from free-text pelvic ultrasound reports. Methods: In this retrospective study, consecutive patients with ovarian masses who underwent pelvic ultrasound were included. Eight LLMs were tested with three reasoning strategies: implicit-knowledge end-to-end, rule-informed end-to-end, and a feature-based hybrid architecture that decoupled feature extraction from rule-based classification. The reference standard was O-RADS categorization established by expert consensus. Results: A total of 310 women with 390 ovarian masses were evaluated. The feature-based hybrid architecture using Gemini 3.6 Flash demonstrated the best performance, achieving an accuracy of 99.2% (387 of 390) and almost perfect agreement with the reference standard (weighted kappa = 1.00; 95% CI: 0.99-1.00). Its performance surpassed that of original clinical reports (accuracy, 87.7% [342 of 390]; weighted kappa = 0.94; 95% CI: 0.91-0.96) and end-to-end LLM strategies (accuracy range, 65.6% [256 of 390] to 95.9% [374 of 390]). For structured feature extraction, Gemini 3.6 Flash demonstrated higher overall accuracy than Claude Fable 5 (98.9% vs 97.8%; P<0.001). The hybrid architecture reduced misclassification errors and mitigated the overstaging tendency observed in original reports. Conclusion: The feature-based hybrid LLM architecture that separates clinical feature extraction from deterministic guideline execution enables highly accurate, reliable, and interpretable automated O-RADS classification, providing a promising approach for standardized, guideline-based clinical decision-making.

View source

Similar papers

Review Open access Aug 2026

Artificial intelligence in clinical decision-making: a comparison of ChatGPT 5.0 and Gemini 3.0 in otologic cases

While both models show potential as clinical decision-support tools, Gemini 3.0 exhibited superior diagnostic performance in real-world otologic cases in the first study evaluating LLMs using real-world otologic data.

Bilge Tuna, Gokhan Tuzemen, Hasan Mutlu · 0 citations
Open access Aug 2026

Comparative assessment of DeepSeek-R1 and GPT-4 for structured ultrasound reporting of adnexal masses

Objective This study primarily evaluated the ability of two large language models (DeepSeek-R1 and GPT-4) to generate structured ultrasound reports from free-text adnexal mass reports. Secondarily, we assessed their accuracy in O-RADS classification and management recommendations, with an exploratory analysis of their...

Meng-Juan Zhang, Bin Wang, Ran Li et al. · 0 citations
Open access Sep 2026

Diagnostic performance and interobserver agreement of O-RADS US v2022 and its combination with serum biomarkers in surgically treated epithelial ovarian-adnexal masses

Objective To evaluate the diagnostic performance of the American College of Radiology (ACR) Ovarian-Adnexal Reporting and Data System for Ultrasound (O-RADS US v2022) and a combined O-RADS US v2022–CA125–HE4 model in differentiating benign from borderline or malignant epithelial ovarian-adnexal masses, and to assess in...

Ying Li, Xiao-Qin He, Li-Ping Chen et al. · 0 citations
Open access Aug 2026

Impact of structured reporting on interpretability of PSMA PET/CT in prostate cancer staging—towards standardised clinical implementation

This study assessed the impact of structured reporting (SR) compared with conventional free-text reporting (FTR) on report quality for PSMA PET/CT for diagnosis and staging of prostate cancer (PC) using a dedicated template. Fifty consecutive patients were included. Original clinical FTRs were compared with SRs retrosp...

Gloria Biechele, M. Schnitzer, Romy Göbel et al. · 0 citations
#large language models Review Open access Sep 2026

RadCoT: a Radiological Chain-of-Thought framework for enhanced error detection in radiology reports

Errors in radiology reports are a major patient-safety concern and are difficult to detect with manual quality assurance (QA). Large language models (LLMs) can assist, but generic prompting does not reflect radiologists’ structured, section-based workflows. To develop and evaluate RadCoT (Radiological Chain-of-Thought)...

Jia Li, Zi-Chun Zhou, Yan-Tao Niu et al. · 0 citations
Open access Sep 2026

Laboratory Medicine Decision Support—Beyond Exam Passing: A Blinded 100-Case Text-Based Benchmark of Diagnostic Accuracy, Management Quality, and Safety for ChatGPT, Gemini, and DeepSeek—LLM Decision Support in Laboratory Medicine

ChatGPT-5.2 had the highest observed performance across several predefined outcomes, although absolute differences were modest for some measures, particularly MCQ accuracy.

K. Ulutaş, A. Pekmezci · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.