Skip to content
Open access

Machine learning models for predicting liver cancer: a real-world cohort study in China

Sep 2026 · Frontiers in Cell and Developmental Biology · Vol 14 · 0 citations · 31 references
Medicine

TL;DR

The findings support the feasibility of leveraging large-scale inpatient laboratory data for risk-stratification model development and show good discrimination and interpretability for liver cancer prediction in a hospitalized real-world cohort using routinely available clinical and laboratory data.

Abstract

Background Liver cancer has high incidence and mortality worldwide, and timely identification is important for improving prognosis. However, prediction models based on routinely available clinical data remain insufficiently evaluated in hospitalized real-world populations. This study aimed to develop and interpret a machine-learning model for liver cancer prediction using multidimensional clinical and laboratory features. Methods This retrospective real-world cohort included 9,284 hospitalized patients from Hainan General Hospital between January 2020 and December 2023. Patients were divided into a training set (n = 6,426) and an internal validation set (n = 2,858). Seventy-four demographic, clinical, and laboratory variables were collected. Feature selection was performed using least absolute shrinkage and selection operator regression in the training set. Six models were developed: logistic regression, random forest, support vector machine, k-nearest neighbors, extreme gradient boosting (XGBoost), and Elastic Net. Discrimination was assessed primarily using the area under the receiver operating characteristic curve (AUC). Decision curve analysis evaluated clinical net benefit, and SHapley Additive exPlanations (SHAP) were used for model interpretation. Results Liver cancer accounted for 48.7% of the training cohort and 48.8% of the validation cohort. LASSO retained 57 nonzero model terms at the minimum-error penalty. XGBoost achieved the highest AUC among all models, with an AUC of 0.909 (95% CI, 0.903–0.915) in the training set and 0.802 (95% CI, 0.788–0.817) in the validation set. At a threshold of 0.5, validation accuracy, sensitivity, specificity, and F1 score were 0.730, 0.758, 0.698, and 0.753, respectively. Decision curve analysis showed that XGBoost provided the greatest net clinical benefit across a broad range of threshold probabilities. SHAP analysis identified serum sialic acid, the aspartate aminotransferase-to-alanine aminotransferase ratio, alkaline phosphatase, basophil percentage, and monocyte-to-lymphocyte ratio as the leading predictors. Conclusion XGBoost showed good discrimination and interpretability for liver cancer prediction in a hospitalized real-world cohort using routinely available clinical and laboratory data. The findings support the feasibility of leveraging large-scale inpatient laboratory data for risk-stratification model development. Serum sialic acid, the aspartate aminotransferase-to-alanine aminotransferase ratio, and alkaline phosphatase were important predictors. Prospective validation in outpatient, first-visit, high-risk, and external populations is required before broader clinical or screening application.

Read PDF

Similar papers

Open access Sep 2026

Research on prognostic prediction of colorectal cancer based on multi-dimensional biomarker features and machine learning models

Background The traditional TNM staging system for colorectal cancer (CRC) prognosis has significant limitations, with prognostic differences exceeding 40% among patients with the same stage. This study aims to construct and validate a machine learning prognostic prediction model based on routinely available clinical, l...

Xiao-Yang Zhang, Xiang-Yong Li, Bo Chen et al. · 0 citations
Open access 2026

An interpretable machine learning model for diagnostic classification of liver cancer using multivariable clinical data.

In this retrospective analysis, we used routinely collected clinical variables to develop a machine learning model for liver cancer diagnosis and examined the variables that contributed to model performance. We studied 3,629 people who were assessed because they were suspected to have liver cancer or other liver diseas...

Dianyu Wang, Xiao Li, Zu-Heng Wang et al. · 0 citations
Open access Aug 2026

Machine learning based survival prediction for liver cancer using cancer registry linked study from the cancer public library database

Prognostic factors differ between young and older patients with liver cancer, but research on predictive models is lacking. We aimed to develop mortality prediction models for younger and older patients with liver cancer and identify key variables. We included data on 4,510 patients diagnosed with liver cancer (2014–20...

H. Kang, W. Jang, Kwang Sun Ryu · 0 citations
Open access Sep 2026

Identification and validation of a chronic kidney disease prediction model after radical nephrectomy: a multicenter cohort retrospective study.

INTRODUCTION Long-term renal function is an important follow-up after radical nephrectomy for renal carcinoma. Predictive models can help identify patients at risk for chronic kidney disease (CKD) progression and enable early intervention. METHODS We retrospectively analyzed 649 patients who underwent radical nephrec...

Xin-Ning Wang, Tian-Wei Zhang, Jing-Xian Wang et al. · 0 citations
Open access Jan 2026

Integrating Serum Lipid Biomarkers Into Machine Learning for the Differential Diagnosis of Breast Nodules

Objective To develop and validate an interpretable machine learning (ML) model for predicting malignant risk in patients with breast nodules using serum lipid biomarkers. Methods This retrospective study included 899 patients with breast nodules (236 malignant) admitted between March 2022 and December 2024. Patients we...

Longmei Chen, Yuzhen Du, Wan-Chao Liu · 0 citations
Sep 2026

Construction and evaluation of a machine-learning-based prediction model for pneumonia in patients with acute leukemia.

An interpretable XGBoost model that accurately predicts pneumonia risk in AL patients based on routine admission data is developed and validated and provides actionable risk stratification to inform preemptive diagnostic and therapeutic strategies.

W. Zhuang, Chen Huang, Xu-Dong Ma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.