BACKGROUND
Cardiovascular disease (CVD) is a leading cause of death worldwide, making early risk prediction essential for improving outcomes. Although artificial intelligence (AI) models promise to improve predictions, questions remain about interpretability, the reliability of risk factors, and the need for cross-validation. This scoping review examined the extent, types, and reporting quality of studies that used AI-based prognostic models to predict CVD risk in individuals without established CVD.
METHODS
A systematic search of PubMed, IEEE Xplore, Web of Science, Scopus, and Google Scholar identified 6,710 records. Screening and data extraction followed the PRISMA-ScR framework and JBI Guidelines for Scoping Reviews. Included studies were evaluated against the TRIPOD-AI reporting checklist. The review protocol was registered with the Open Science Framework (OSF: https://osf.io/8nq6s/).
RESULTS
Thirty studies met the inclusion criteria; all published after 2017. Most studies utilized existing models on large datasets, predominantly leveraging unimodal clinical data and established machine learning algorithms such as Random Forests and Support Vector Machines. Twenty-four of the 30 studies used unimodal approaches, and the six multimodal studies demonstrated consistently strong performance, but they rest on a small, heterogeneous set of studies. Twelve studies conducted direct comparisons with traditional risk scores such as the Framingham Risk Score, showing comparable or modestly improved discrimination, although methodological heterogeneity limits the strength of these conclusions. External validation was reported in only seven studies, and calibration (agreement between the predicted and observed event rates) was not reported in any of the 30 included studies. Sensitivity, which determines a model's ability to identify truly high-risk individuals, was the least reported metric, appearing in only four studies, and no study reported a decision curve analysis.
CONCLUSION
AI-based CVD risk prediction tools show promise but have critical gaps in validation, calibration reporting, and clinical utility assessment, which currently preclude clinical deployment. Future research should prioritize external validation across diverse populations, mandatory reporting of calibration and sensitivity alongside discrimination metrics, as well as the adoption of reporting standards and appraisal tools such as Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis + Artificial Intelligence (TRIPOD + AI) or Prediction Model Risk of Bias Assessment Tool + Artificial Intelligence (PROBAST + AI) to improve transparency and reproducibility.
S. Pai, Rekha Subramanian, Lakshmi Krishnan et al.· BMC Medical Informatics and...· 0 citations
Oral cancer is a leading cause of mortality in low-to-middle-income countries, where a shortage of specialists delays diagnosis. While point-of-care screening via smartphones offers a scalable solution, developing robust AI for resource-constrained settings poses significant challenges, including class imbalance in training data, variable data quality, and computational constraints on edge devices. In this paper, we present the optimisation of lightweight deep learning models for smartphone-based oral cancer screening. Using a diverse, multi-centre retrospective dataset of approximately 30,000 images acquired over a decade, we systematically evaluate state-of-the-art convolutional, transformer, and hybrid architectures. Through rigorous pipeline ablation, we demonstrate that directly optimising hybrid architectures for the edge strictly outperforms computationally heavy paradigms, such as large models or knowledge distillation. Furthermore, interpretability analysis and simulated noise-stress tests revealed that the system anchors on clinical features and remains robust to unstructured sensor noise, despite vulnerabilities to impulse bit errors. In the held-out test set, our optimised MobileViTv2 models achieved an average sensitivity of 83.2 $\pm$ 1.5% and an average specificity of 86.0 $\pm$ 0.8%, with the best model exhibiting 87.4% sensitivity, 86.5% specificity, and a critical negative predictive value of 97.2% with reference to specialist labels. These results confirm that with targeted architectural selection and streamlined optimisation, interpretable and robust lightweight AI models exhibit high potential for edge deployment to enable automated triage in primary care settings.
Siddhant Bharadwaj, Aakash Shedsale, T. Subramanya et al.· 0 citations