Jul 2026· European Heart Journal - Digital Health· Vol 7· 0 citations· 27 references
Medicine
TL;DR
P Pipelines combining expert-curated knowledge injection, LLM-based clinical data extraction, and deterministic score calculation enable accurate and scalable cardiovascular risk score computation from unstructured real-world clinical data, outperforming physician-documented scores.
Abstract
Abstract Aims Risk scores are essential to evidence-based cardiovascular care, but manual calculation is labour intensive and error prone. Large language models (LLMs) could automate this process, yet LLMs are limited by their propensity for calculation errors and factual hallucinations. Pipelines separating LLM-based data extraction from deterministic score computation may improve reliability and transparency. Methods and results We conducted a retrospective diagnostic study at a quaternary heart centre in Germany (January 2020 to July 2023). Patients with atrial fibrillation (n = 179) from an ablation registry and patients with severe aortic stenosis (n = 76) evaluated by a heart team were included. Six LLMs (GPT-5.2, Gemini 3.1 Pro, DeepSeek-R1, Qwen3, GPT-OSS 120B, and Kimi K2.5) were tested in standalone, retrieval-augmented generation (RAG), and pipeline configurations to compute HAS-BLED, CHA2DS2-VASc, and EuroSCORE II scores from routine clinical reports. Accuracy was assessed against expert-adjudicated ground truth using root mean squared error (RMSE) and Krippendorff’s α to evaluate numerical deviation and categorical agreement, respectively. Pipeline-generated scores showed substantially higher agreement with expert adjudication than standalone LLMs, LLMs with RAG, and treating physicians (mean Krippendorff’s α: 0.78 vs. 0.32 vs. 0.39 vs. 0.31) and lower deviation from ground truth (mean RMSE: 0.89 vs. 5.81 vs. 1.85 vs. 1.34). Conclusion Pipelines combining expert-curated knowledge injection, LLM-based clinical data extraction, and deterministic score calculation enable accurate and scalable cardiovascular risk score computation from unstructured real-world clinical data, outperforming physician-documented scores. Such pipelines could form the basis for clinical decision-support systems that automate routine risk assessment, reduce clinician workload, and promote more consistent evidence-based care.
Background: Large language models (LLMs) are increasingly used in patient-facing clinical applications, creating a need for scalable methods to evaluate the accuracy, safety, and appropriateness of their responses. Although physician review remains the reference standard, it is resource-intensive and difficult to scale...
B. Collaço, Nadia G. Wood, Yun-Guo Yu et al.· Bioengineering· 0 citations
Highlights
Large Language Models (LLMs) enable automation of initial abstract screening in systematic reviews, significantly reducing manual workload.
The effectiveness of LLMs heavily depends on prompt engineering, which must clearly translate inclusion and exclusion criteria into actionable instructio...
Shatskiy Alexander S., Ehab M. Deigheidy, S. E. Masyutina et al.· Complex Issues of Cardiovasc...· 0 citations
Risk stratification in heart failure (HF) supports clinical decisions, yet existing tools face adoption barriers: conventional scores show modest discrimination and depend on specialised tests (i.e., echocardiography), while artificial intelligence (AI) models require rich longitudinal data and infrastructure. Both a...
N. Ahmed, N. Conrad, M. Wamil et al.· npj Digital Medicine· 0 citations
Manual clinical DNA variant classification is the bottleneck of every clinical and research rare disease workflow. The process typically requires a curator to assemble evidence from numerous databases, weigh 28 criteria, reconcile competing evidence, and produce a defensible case for the final classification. Additiona...
Jamie-Lee M. Thompson, Debjani Das, Sally L. Dunwoodie et al.· bioRxiv· 0 citations
Statistical analysis of clinical data requires expertise in medical statistics. Large language models (LLMs) are increasingly used for code generation and may support both descriptive and advanced analyses, but their reliability remains uncertain. This study evaluated whether five current LLMs (GPT 5.3, Claude Sonnet 4...
J. Sam, T. Spreuer, M. Berger et al.· Studies in Health Technology...· 0 citations
Background: Large language models (LLMs) are increasingly accessible to healthcare providers and patients for clinical decision support, yet their ability to perform precise mathematical calculations required for validated risk scores remains unexplored, and errors could compromise patient safety. The EuroSCORE II requ...
Philip DiCicco, W. Bennar, Corentin Volet et al.· Cardiovascular Medicine· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.