Skip to content
Open access

AI-ECG Risk Stratification for Atrial Fibrillation

Jul 2026 · JACC: Advances · Vol 5 · 0 citations · 11 references
Medicine

TL;DR

AI-ECG provides rapid, low-cost AF risk estimation from a single sinus rhythm ECG; however, its predictive performance is modest compared with clinical scores.

Abstract

Background Artificial intelligence–enabled electrocardiography (AI-ECG) has emerged as a potential method for identifying atrial fibrillation (AF) from sinus rhythm. However, its clinical utility and interpretability in routine practice remain uncertain. Objectives The objective of the study was to assess the performance and explainability of an AI-integrated ECG system for AF risk stratification in a multicenter real-world cohort. Methods We enrolled 665 patients aged ≥40 years who underwent 12-lead ECGs using an AI-enabled electrocardiograph (FCP-9900). The device automatically assigned AF risk into 4 categories (low, mid-low, mid-high, and high). Machine-learning models—support vector machine, adaptive boosting, and artificial neural networks—were developed using clinical and ECG-derived variables. Internal validation used stratified 10-fold cross-validation, and external validation used an independent cohort. Feature contributions were assessed with SHapley Additive exPlanations. Results AF prevalence increased across AI-ECG risk categories, with significantly higher odds in the mid-high and high groups vs low. Model 2, which incorporated CHADS2 and CHA2DS2-VASc scores, achieved strong discrimination in internal and external validation (support vector machine AUC 1.00; adaptive boosting 0.97-0.98; artificial neural network 0.89-0.95), outperforming AI-ECG alone (area under the receiver operating characteristic curve: 0.64-0.69). SHapley Additive exPlanations analysis showed CHA2DS2-VASc as the most influential predictor, whereas AI-ECG provided modest incremental value. Conclusions AI-ECG provides rapid, low-cost AF risk estimation from a single sinus rhythm ECG; however, its predictive performance is modest compared with clinical scores. At present, AI-ECG may complement, but not replace, traditional risk stratification.

Read PDF

Similar papers

Review Open access Jul 2026

Impact of an AI algorithm for multi-day prediction of incident atrial fibrillation on clinical decision-making: PROVISION-AF study.

In conclusion, AI-derived risk estimates improved physician risk discrimination in a structured simulated survey, particularly in non-specialist settings, supporting their potential role as a digital decision-support tool.

Yeji Kim, Bogeun Kim, J. Yoon et al. · 0 citations
Review Open access Aug 2026

Artificial intelligence–enabled electrocardiography for assessment of left ventricular systolic dysfunction in the era of foundation models

Artificial intelligence (AI) applied to the standard 12-lead electrocardiogram (AI-ECG) is being developed as a scalable approach to screen for left ventricular systolic dysfunction (LVSD) and support triage for confirmatory testing. Supervised models trained on paired ECG–echocardiography data show high discrimination for reduced ejection fraction across thresholds and can identify individuals at higher risk of subsequent LV dysfunction despite a normal baseline echocardiogram. External validation of an FDA-cleared ECG-AI device across four geographically diverse U.S. health systems confirmed strong diagnostic accuracy, though signal-format compatibility and quality gating meaningfully affect real-world yield. Two pragmatic randomized trials demonstrate practice-level impact. In primary care, AI-ECG increased the number of new low-ejection-fraction diagnoses and directed echocardiography preferentially to screen-positive patients. In non-cardiology inpatient wards, AI alerts improved diagnostic yield through increased cardiology consultation rather than increased imaging volume. In emergency-department patients with dyspnea, AI-ECG supports a prioritization role with high negative predictive value, outperforming NT-proBNP, but requires confirmatory imaging given prevalence-dependent positive predictive value. In population cohorts, adding AI-ECG signals to PREVENT-HF improves near-term heart-failure risk discrimination and reclassification, though without demonstrated benefit on clinical outcomes such as heart-failure hospitalization or mortality. Foundation models pretrained on large ECG datasets reduce labeled-data requirements and improve transportability, but prospective echocardiography-anchored validation is required before broader deployment. FDA-cleared software is available for left ventricular ejection fraction ≤40% screening from 12-lead ECGs as clinician decision support. This review summarizes performance across thresholds and care settings, outlines threshold selection and calibration, and defines priorities for outcome-oriented trials, equitable deployment, and implementation governance.

A. Bollmann, V. Pradler, D. Husser et al. · 0 citations
Open access Aug 2026

Using AI-ECG to Stratify Long-Term Mortality Risk and Prognosis in TAVR Patients.

BACKGROUND Long-term mortality remains unsatisfactorily high after transcatheter aortic valve replacement (TAVR). Conventional risk models are limited in capturing subclinical electrophysiological alterations associated with poor prognosis, which can be identified on routine preoperative electrocardiograms. OBJECTIVES The authors aim to develop and validate an artificial intelligence-enhanced electrocardiogram (AI-ECG) model for predicting long-term mortality in post-TAVR patients. METHODS A total of 711 patients with severe aortic stenosis undergoing TAVR were enrolled from 2 centers. Patients from one center were divided into training and internal validation sets (7:3), and participants from another center served as the external validation cohort. Preoperative electrocardiogram images were analyzed using a Residual Network-18 model to generate mortality risk stratification. The primary endpoint was 3-year all-cause death. RESULTS The AI-ECG model demonstrated comparable discrimination between the internal and external patient cohorts, with areas under the receiver operating characteristics curve of 0.767 (95% CI: 0.657-0.877) vs 0.712 (95% CI: 0.627-0.795) (P for DeLong test = 0.428). High-risk patients (15.5% [39 of 251]) exhibited a 61.5% (24 of 39, 95% CI: 42.8%-74.1%) 3-year mortality rate vs 16.5% (35 of 212, 95% CI: 11.4%-21.4%) in low-risk patients (84.5% [212 of 251]) (log-rank P < 0.001). Adjusted for comorbidities, high-risk classification independently predicted mortality (adjusted HR: 3.49; 95% CI: 1.96-6.22). Subgroup analysis did not reveal significant interaction effects of the AI-ECG model across different patient populations. Decision curve analysis confirmed clinical net benefit across threshold probabilities (0.05-0.60). CONCLUSIONS The AI-ECG model provides noninvasive and accurate long-term risk stratification for TAVR patients, with promising clinical application value for individualized follow-up management.

Wence Shi, Peirou Yan, Qifeng Zhu et al. · 0 citations
Open access Jul 2026

Artificial intelligence-based ECG reconstruction error as a continuous predictor of all-cause mortality: a multi-cohort retrospective validation study

Background Recent artificial intelligence (AI) models applied to the electrocardiogram (ECG) for risk stratification typically rely on supervised learning, defining risk as the error relative to an external target such as age or sex. This couples the risk score to the choice of target rather than the cardiac signal alone, and may limit generalisability. We aimed to develop a self-supervised AI-ECG risk score based on the error in reconstructing a partially masked ECG. Methods A transformer-based masked autoencoder was trained on 85% of the CODE dataset (n = 7,212,109 ECGs) to reconstruct ECG signals from partially masked inputs. The association between reconstruction error and all-cause mortality was assessed internally in CODE-15% and externally validated in four independent cohorts: MIMIC-IV-ECG (critical care, US), HEEDB (hospital, US), CHRIS (population-based, Italy), and Innsbruck (cardiology centre, Austria). A binary risk score (>1 SD above the CODE-15% mean) was additionally evaluated in these cohorts and in the UK Biobank (population-based, UK). Findings In Cox proportional hazards models adjusted for age and sex, each 1-SD increase in reconstruction error was associated with higher all-cause mortality (all p<0.001; cohort median follow-up 1.4-11.0 years): CODE-15% (HR 1.39, 95% CI 1.37-1.42), MIMIC-IV-ECG (HR 1.39, 95% CI 1.37-1.40), HEEDB (HR 1.41, 95% CI 1.40-1.41), Innsbruck (HR 1.23, 95% CI 1.21-1.26), and CHRIS (HR 1.25, 95% CI 1.14-1.38). The binary threshold identified a high-risk group with increased mortality in all six cohorts, including the UK Biobank (HR 1.27, 95% CI 1.08-1.50, p=0.004). Interpretation Reconstruction error is a generalisable predictor of all-cause mortality across diverse clinical and population-based settings. Unlike supervised approaches, it reflects the model's uncertainty about the ECG signal itself rather than error relative to an external target, providing a direct measure of how much each recording deviates from normal cardiac electrical patterns.

A. Nicolson, S. Pröll, R. Lunelli et al. · 0 citations
Open access Feb 2026

Artificial intelligence shows comparable or improved performance to traditional risk models in predicting atrial fibrillation after cryptogenic stroke

In patients with cryptogenic stroke receiving an ICM, the ECG-AI score showed modest discrimination for AF detection, outperforming CHA2DS2-VA and HAVOC, but not Brown ESUS-AF, which indicates a possible role for AI-driven ECG analysis in risk stratification.

F. Wouters, M. Barthels, J. Vranken et al. · 0 citations
Review Open access Aug 2026

Artificial intelligence in electrocardiogram interpretation for cardiovascular diagnosis and risk prediction: a systematic review of evidence through May 2026

Artificial intelligence-enabled electrocardiogram (ECG) analysis has expanded across cardiovascular diagnosis, monitoring, and risk prediction, but published studies vary in methodology, validation, and clinical applicability. We systematically reviewed AI applications in ECG interpretation for cardiovascular diagnosis and risk prediction to define the evidence available through May 2026. This systematic review followed PRISMA 2020. PubMed and the Cochrane Library were searched from database inception through May 2026 for original human studies evaluating artificial intelligence, machine learning, or deep learning applications in ECG-based cardiovascular detection, classification, diagnosis, monitoring, or risk prediction. Risk of bias was assessed using PROBAST. Because of substantial heterogeneity in study populations, ECG modalities, model architectures, validation strategies, and reported outcomes, findings were synthesized narratively. A total of 108 studies were included. Arrhythmia detection, particularly atrial fibrillation (AF), was the most mature and frequently evaluated application domain. Benchmark-based studies commonly reported high apparent performance, whereas larger clinical cohorts and externally validated studies generally showed more moderate but more clinically credible estimates. The most common methodological limitations were concentrated in the analysis domain, particularly limited external validation, overfitting risk, unclear train-test separation, class imbalance, and incomplete reporting of calibration and robustness. AI-enabled ECG interpretation shows strongest support for arrhythmia detection and automated ECG classification, while structural disease screening and prognostic modeling remain promising but less mature. Future studies should prioritize prospective, multicenter, externally validated, and workflow-integrated designs with transparent reporting, calibration assessment, cost-effectiveness evaluation, and equitable assessment across diverse populations.

Mohamed O Elhussain, Esra M Abdalla, Ragda Ali et al. · 0 citations