Background: Heart failure with reduced ejection fraction (HFrEF) remains a major global health burden. Most electrocardiogram (ECG)-based artificial intelligence models are limited to diagnostic tasks or fixed-horizon prognostic classification and provide little insight into the temporal evolution of risk. In addition, concerns regarding model interpretability continue to impede clinical adoption. Whether deep learning applied to ECGs can deliver individualized, time-resolved, and biologically interpretable risk estimates for incident HFrEF across diverse populations remains uncertain. Methods: We developed a convolutional neural network-based survival model using raw 12-lead ECGs from Zhongshan Hospital (SHZS) and externally validated it in independent cohorts from Shanghai Tenth People's Hospital (SHTP) and Beth Israel Deaconess Medical Center (BIDMC). The model generated individualized, day-by-day probabilities of incident HFrEF over a 5-year horizon. Performance was comprehensively evaluated using discrimination, calibration, precision-recall characteristics, clinical utility, and risk stratification metrics, with subgroup analyses across age, sex, and race to assess generalizability. Model interpretability was examined using complementary representation and attention-based frameworks. Results: In 458,884 patients, the survival model demonstrated strong and stable discrimination across cohorts, with overall C-indices of 0.971 (95% CI, 0.965-0.976) in SHZS, 0.945 (95% CI, 0.938-0.950) in SHTP, and 0.855 (95% CI, 0.850-0.860) in BIDMC, and consistently high time-dependent AUROC values across the 1-5-year horizons. Calibration showed close agreement between predicted and observed risks, and decision curve analyses indicated meaningful net clinical benefit across a broad range of thresholds. Kaplan-Meier curves showed clear stratification across predicted risk groups. Interpretability analyses identified physiologically coherent ECG features related to QRS duration, heart rate, and QT interval that were associated with predicted risk. Conclusion: This ECG-based deep learning survival model provides individualized, time-resolved, and clinically interpretable estimates of future HFrEF risk with robust performance across multinational cohorts. These findings support the potential of AI-enabled ECG analysis as an accessible tool for early HFrEF risk stratification within routine clinical workflows.
L. Pan, S. Li, J. Huo et al.· medRxiv· 0 citations
Background: Regurgitant valvular heart disease (rVHD) is a major cause of cardiovascular morbidity. Echocardiography is the diagnostic standard but is resource-intensive for large-scale screening. Electrocardiography (ECG) has shown promise for predicting incident rVHD, yet performance varies across phenotypes, particularly for aortic regurgitation (AR). Chest radiography (CXR) provides complementary structural and hemodynamic information. We hypothesized that a multimodal model integrating ECG and CXR would improve prediction of incident moderate-to-severe rVHD. Methods: In this retrospective multicenter study, we identified 212,888 paired ECG-CXR examinations from 116,380 patients across two Chinese centers. Baseline ECG and CXR were obtained within 60 days of echocardiography. Outcome was progression to moderate-to-severe AR, mitral regurgitation (MR), or tricuspid regurgitation (TR). We developed a multimodal neural network with pretrained unimodal encoders, token-level cross-modal fusion, and a class-specific gating mechanism that adaptively weighted ECG-only, CXR-only, and fused predictions. Performance was assessed using C-index, AUROC, AUPRC, decision curve analysis, net reclassification improvement (NRI), and Kaplan-Meier stratification. Results: Multimodal fusion consistently outperformed unimodal models across all phenotypes. For AR, C-index improved from 0.616 (ECG-only) to 0.713 (multimodal; AUROC 0.729, AUPRC 0.972). For MR, multimodal C-index was 0.801 (AUROC 0.814, AUPRC 0.972), versus 0.782 for ECG and 0.775 for CXR alone. For TR, multimodal and CXR-only models showed similar discrimination (C-index 0.802), but multimodal fusion yielded greater net benefit on decision curve analysis. NRI was positive across all time horizons (1-5 years) for all valve types. Grad-CAM interpretability analyses revealed that ECG attention localized to leads II, V-V (AR), leads I, II, aVF, V-V (MR), and inferior/right precordial leads (TR); CXR attention highlighted chamber-specific enlargement and pulmonary congestion patterns consistent with pathophysiology. Conclusion: A multimodal deep learning model integrating ECG and CXR significantly improved prediction of incident rVHD compared with ECG alone, with the greatest benefit observed for AR. The model leveraged complementary electrical and structural information, demonstrated biological plausibility through interpretability analyses, and provided consistent clinical utility. Given the widespread availability and low cost of both modalities, this approach offers a scalable tool for risk stratification in routine care. Prospective studies are warranted to validate clinical implementation.
S. Li, B. Zhang, L. Pan et al.· medRxiv· 0 citations