Skip to content
Review Open access

Applications and limitations of machine learning in clinical biostatistics: a narrative review

Jul 2026 · Journal of Clinical Technology and Theory · 0 citations

Abstract

Clinical biostatistics has traditionally relied on regression-based models to explain associations, estimate risks, and support medical decision-making. The growth of electronic health records, imaging data, omics data, and follow-up data has made clinical information larger, less tidy, and more difficult to model with only classical methods. Machine learning is increasingly used in this setting because it can capture non-linear patterns and handle high-dimensional predictors. This narrative review summarizes applications of supervised learning, unsupervised learning, and deep learning in clinical diagnosis, prognosis prediction, patient stratification, and biomarker discovery. Literature was selected from PubMed, Web of Science, and Google Scholar, with emphasis on studies and reporting guidelines published from 2019 to 2025. This review finds that machine learning is useful when the clinical question is clear, data quality is acceptable, and validation is strict. However, a stronger algorithm does not automatically become a better clinical tool. Main limitations include weak interpretability, biased training data, poor transportability, overreliance on Area Under the Curve (AUC), and incomplete reporting. Machine learning should therefore be treated as a complement to traditional biostatistics rather than a simple replacement.

Read PDF