Skip to content
Open access

Development of an LLM pipeline exceeding physician-documented cardiovascular risk scores under routine clinical conditions

Jul 2026 · European Heart Journal - Digital Health · Vol 7 · 0 citations · 27 references
Medicine

TL;DR

P Pipelines combining expert-curated knowledge injection, LLM-based clinical data extraction, and deterministic score calculation enable accurate and scalable cardiovascular risk score computation from unstructured real-world clinical data, outperforming physician-documented scores.

Abstract

Abstract Aims Risk scores are essential to evidence-based cardiovascular care, but manual calculation is labour intensive and error prone. Large language models (LLMs) could automate this process, yet LLMs are limited by their propensity for calculation errors and factual hallucinations. Pipelines separating LLM-based data extraction from deterministic score computation may improve reliability and transparency. Methods and results We conducted a retrospective diagnostic study at a quaternary heart centre in Germany (January 2020 to July 2023). Patients with atrial fibrillation (n = 179) from an ablation registry and patients with severe aortic stenosis (n = 76) evaluated by a heart team were included. Six LLMs (GPT-5.2, Gemini 3.1 Pro, DeepSeek-R1, Qwen3, GPT-OSS 120B, and Kimi K2.5) were tested in standalone, retrieval-augmented generation (RAG), and pipeline configurations to compute HAS-BLED, CHA2DS2-VASc, and EuroSCORE II scores from routine clinical reports. Accuracy was assessed against expert-adjudicated ground truth using root mean squared error (RMSE) and Krippendorff’s α to evaluate numerical deviation and categorical agreement, respectively. Pipeline-generated scores showed substantially higher agreement with expert adjudication than standalone LLMs, LLMs with RAG, and treating physicians (mean Krippendorff’s α: 0.78 vs. 0.32 vs. 0.39 vs. 0.31) and lower deviation from ground truth (mean RMSE: 0.89 vs. 5.81 vs. 1.85 vs. 1.34). Conclusion Pipelines combining expert-curated knowledge injection, LLM-based clinical data extraction, and deterministic score calculation enable accurate and scalable cardiovascular risk score computation from unstructured real-world clinical data, outperforming physician-documented scores. Such pipelines could form the basis for clinical decision-support systems that automate routine risk assessment, reduce clinician workload, and promote more consistent evidence-based care.

Read PDF

Similar papers

Review Open access Sep 2026

Development and Clinical Validation of an Automated LLM Judge for Evaluating Perioperative Patient Questions

Background: Large language models (LLMs) are increasingly used in patient-facing clinical applications, creating a need for scalable methods to evaluate the accuracy, safety, and appropriateness of their responses. Although physician review remains the reference standard, it is resource-intensive and difficult to scale...

B. Collaço, Nadia G. Wood, Yun-Guo Yu et al. · 0 citations
Review Open access Sep 2026

USING LARGE LANGUAGE MODELS FOR LITERATURE SEARCH IN CARDIOVASCULAR SURGERY SYSTEMATIC REVIEWS AND META-ANALYSES

Highlights Large Language Models (LLMs) enable automation of initial abstract screening in systematic reviews, significantly reducing manual workload. The effectiveness of LLMs heavily depends on prompt engineering, which must clearly translate inclusion and exclusion criteria into actionable instructio...

Shatskiy Alexander S., Ehab M. Deigheidy, S. E. Masyutina et al. · 0 citations
Open access Sep 2026

Development and validation of a parsimonious AI-based mortality risk score for heart failure

Risk stratification in heart failure (HF) supports clinical decisions, yet existing tools face adoption barriers: conventional scores show modest discrimination and depend on specialised tests (i.e., echocardiography), while artificial intelligence (AI) models require rich longitudinal data and infrastructure. Both a...

N. Ahmed, N. Conrad, M. Wamil et al. · 0 citations
Open access Sep 2026

HeartVar: An LLM-Assisted Tool for Clinical Classification of Variants in Cardiovascular Disease Cohorts

Manual clinical DNA variant classification is the bottleneck of every clinical and research rare disease workflow. The process typically requires a curator to assemble evidence from numerous databases, weigh 28 criteria, reconcile competing evidence, and produce a defensible case for the final classification. Additiona...

Jamie-Lee M. Thompson, Debjani Das, Sally L. Dunwoodie et al. · 0 citations
#large language models Open access Sep 2026

Benchmarking AI Vibe Coding for Clinical Statistical Analysis: A Structured Evaluation in Pulmonary Hypertension Research.

Statistical analysis of clinical data requires expertise in medical statistics. Large language models (LLMs) are increasingly used for code generation and may support both descriptive and advanced analyses, but their reliability remains uncertain. This study evaluated whether five current LLMs (GPT 5.3, Claude Sonnet 4...

J. Sam, T. Spreuer, M. Berger et al. · 0 citations
Open access Aug 2026

Large Language Model Chatbots Cannot Reliably Calculate Clinical Risk Scores—A Comparative Accuracy Study

Background: Large language models (LLMs) are increasingly accessible to healthcare providers and patients for clinical decision support, yet their ability to perform precise mathematical calculations required for validated risk scores remains unexplored, and errors could compromise patient safety. The EuroSCORE II requ...

Philip DiCicco, W. Bennar, Corentin Volet et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.