Identification of Candidate Diagnostic Biomarkers and Construction of an Interpretable Machine-Learning Model for Acute Myeloid Leukemia
Abstract
Acute myeloid leukemia (AML) is a molecularly heterogeneous disease, with considerable variation in molecular abnormalities and gene-expression patterns among patients. This study aimed to identify a compact diagnostic gene signature and develop an interpretable machine-learning model for distinguishing AML from normal samples. Gene-expression data from GSE9476, comprising 26 AML and 38 normal samples, were divided into a 51-sample training set and an independent 13-sample test set. Training-restricted differential analysis identified 283 significant genes. Least absolute shrinkage and selection operator regression, combined with Random Forest screening, reduced this set to ALDH1A1, GPX1, CAPRIN2, MFSD10 and IL3RA. Logistic Regression, linear support vector machine and Random Forest models were compared by five-fold stratified cross-validation. Logistic Regression was retained as the primary model and achieved a test ROC-AUC of 1.000, with 92.3% accuracy, 80.0% sensitivity and 100% specificity. Shapley additive explanations ranked GPX1, ALDH1A1 and IL3RA as the strongest contributors, with effects aligned with the differential expression results. External validation in 77 GSE30029 samples produced a ROC-AUC of 0.969 (95% confidence interval: 0.922–1.000), 92.2% accuracy and 100% specificity. Agreement across independent cohorts and microarray platforms supports the five-gene panel as a reproducible candidate signature. Larger multicenter cohorts and experimental assays remain necessary before clinical application.