Comparative Analysis of Logistic Regression and Random Forest Models for Cardiovascular Disease Prediction Using Clinical Data
Abstract
. Cardiovascular disease remains a significant global health burden, making it crucial to construct accurate risk prediction models. Analyzing a large clinical dataset, this study evaluates the performance of logistic regression and random forest models in predicting cardiovascular and cerebrovascular diseases. Data are divided into training and testing sets, and multiple indicators such as accuracy, precision, recall, F1 score, and area under the ROC curve (AUC) are used to measure the model's effectiveness. Results show that both models exhibited good predictive performance, with slightly better accuracy in random forest classification and outstanding discriminative ability in logistic regression (AUC of 0.78). Studying ensemble models can bring marginal performance improvements, but traditional statistical methods still have significant advantages and interpretability in clinical risk prediction. These conclusions provide clinically actionable guidance for clinicians and researchers to select appropriate machine learning models—such as preferring logistic regression for interpretable risk stratification or random forest for slightly higher accuracy—when screening for cardiovascular and cerebrovascular diseases in clinical practice.