The feature-based hybrid LLM architecture that separates clinical feature extraction from deterministic guideline execution enables highly accurate, reliable, and interpretable automated O-RADS classification, providing a promising approach for standardized, guideline-based clinical decision-making.
Abstract
Background: Automating clinical guideline-based decision-making with large language models (LLMs) remains challenging because of reliability, hallucination, and limited interpretability. We compared the performance of LLMs and reasoning strategies for automated Ovarian-Adnexal Reporting and Data System (O-RADS) classification from free-text pelvic ultrasound reports. Methods: In this retrospective study, consecutive patients with ovarian masses who underwent pelvic ultrasound were included. Eight LLMs were tested with three reasoning strategies: implicit-knowledge end-to-end, rule-informed end-to-end, and a feature-based hybrid architecture that decoupled feature extraction from rule-based classification. The reference standard was O-RADS categorization established by expert consensus. Results: A total of 310 women with 390 ovarian masses were evaluated. The feature-based hybrid architecture using Gemini 3.6 Flash demonstrated the best performance, achieving an accuracy of 99.2% (387 of 390) and almost perfect agreement with the reference standard (weighted kappa = 1.00; 95% CI: 0.99-1.00). Its performance surpassed that of original clinical reports (accuracy, 87.7% [342 of 390]; weighted kappa = 0.94; 95% CI: 0.91-0.96) and end-to-end LLM strategies (accuracy range, 65.6% [256 of 390] to 95.9% [374 of 390]). For structured feature extraction, Gemini 3.6 Flash demonstrated higher overall accuracy than Claude Fable 5 (98.9% vs 97.8%; P<0.001). The hybrid architecture reduced misclassification errors and mitigated the overstaging tendency observed in original reports. Conclusion: The feature-based hybrid LLM architecture that separates clinical feature extraction from deterministic guideline execution enables highly accurate, reliable, and interpretable automated O-RADS classification, providing a promising approach for standardized, guideline-based clinical decision-making.
While both models show potential as clinical decision-support tools, Gemini 3.0 exhibited superior diagnostic performance in real-world otologic cases in the first study evaluating LLMs using real-world otologic data.
Bilge Tuna, Gokhan Tuzemen, Hasan Mutlu· BMC Medical Informatics and...· 0 citations
Objective This study primarily evaluated the ability of two large language models (DeepSeek-R1 and GPT-4) to generate structured ultrasound reports from free-text adnexal mass reports. Secondarily, we assessed their accuracy in O-RADS classification and management recommendations, with an exploratory analysis of their...
Meng-Juan Zhang, Bin Wang, Ran Li et al.· Frontiers in Oncology· 0 citations
Objective To evaluate the diagnostic performance of the American College of Radiology (ACR) Ovarian-Adnexal Reporting and Data System for Ultrasound (O-RADS US v2022) and a combined O-RADS US v2022–CA125–HE4 model in differentiating benign from borderline or malignant epithelial ovarian-adnexal masses, and to assess in...
Ying Li, Xiao-Qin He, Li-Ping Chen et al.· Frontiers in Medicine· 0 citations
This study assessed the impact of structured reporting (SR) compared with conventional free-text reporting (FTR) on report quality for PSMA PET/CT for diagnosis and staging of prostate cancer (PC) using a dedicated template. Fifty consecutive patients were included. Original clinical FTRs were compared with SRs retrosp...
Gloria Biechele, M. Schnitzer, Romy Göbel et al.· European Radiology Experimen...· 0 citations
Errors in radiology reports are a major patient-safety concern and are difficult to detect with manual quality assurance (QA). Large language models (LLMs) can assist, but generic prompting does not reflect radiologists’ structured, section-based workflows. To develop and evaluate RadCoT (Radiological Chain-of-Thought)...
Jia Li, Zi-Chun Zhou, Yan-Tao Niu et al.· European Radiology Experimen...· 0 citations
ChatGPT-5.2 had the highest observed performance across several predefined outcomes, although absolute differences were modest for some measures, particularly MCQ accuracy.
K. Ulutaş, A. Pekmezci· Diagnostics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.