Skip to content
Open access

Large Language Models for Heterogeneous Data Mining in Liver Disease: Framework Development and Retrospective Validation Study

Sep 2026 · Journal of Medical Internet Research · Vol 28, pp. e92921-e92921 · 0 citations · 60 references
Medicine

TL;DR

In this single-center retrospective cohort of patients with clear-cut, nonoverlapping liver disease etiologies, LLM-derived embeddings provided complementary information to structured laboratory variables for multiclass liver disease classification.

Abstract

Abstract Background Differentiating among liver disease entities such as autoimmune liver disease (AILD), drug-induced liver injury (DILI), and chronic hepatitis B (CHB) remains clinically challenging due to overlapping clinical manifestations and nonspecific laboratory findings. Conventional machine learning (ML) approaches rely mainly on structured laboratory data, whereas free-text clinical reports and other heterogeneous electronic medical record data are often underused. Large language models (LLMs) may provide a strategy for encoding heterogeneous clinical information, yet their usefulness for liver disease classification remains insufficiently evaluated. Objective This study aimed to evaluate the usefulness of LLM-derived embeddings for clinical data mining in liver disease and to determine whether integrating these embeddings with laboratory variables improves classification across broad disease categories and closely related subtypes. Methods We retrospectively analyzed electronic medical record data from 7543 patients with nonoverlapping liver disease etiologies treated at Beijing Youan Hospital, Capital Medical University, between 2010 and 2025. Three LLMs (Qwen3, Huatuo-o1, and II-Medical) generated semantic embeddings from standardized clinical text, combining free-text examination reports, and structured clinical observations. Performance was assessed in a 3-class etiological task (AILD, DILI, and CHB) and a 4-class task further subclassifying AILD into autoimmune hepatitis and primary biliary cholangitis. We compared embedding-only models, LLM-integrated ML models, and an ML-only baseline using the same structured variable set and preprocessing pipeline, with lightweight natural language processing encoders and zero-shot LLM reasoning as additional comparators. Models were developed using 5-fold cross-validation and evaluated on an internal holdout set using accuracy, macroaveraged precision, recall, and F1-score. Results In the 3-class task, the LLM-integrated ML models achieved macro F1-scores of 0.835‐0.837, compared with 0.791 for the ML-only baseline, with corresponding accuracies of 0.925‐0.929 versus 0.893. In the 4-class task, the LLM-integrated ML models achieved macro F1-scores of 0.717‐0.734, compared with 0.665 for the ML-only baseline, with corresponding accuracies of 0.920‐0.922 versus 0.874. A temporal split sensitivity analysis using cases from 2010 to 2019 for training and cases from 2020 to 2025 for testing showed that the relative advantage of LLM-integrated ML models over the ML-only baseline was preserved. Direct zero-shot LLM reasoning and lightweight natural language processing encoders performed below the embedding-based integrated models. Conclusions In this single-center retrospective cohort of patients with clear-cut, nonoverlapping liver disease etiologies, LLM-derived embeddings provided complementary information to structured laboratory variables for multiclass liver disease classification. The integrated framework showed improved internal validation performance compared with the ML-only model, particularly for non-CHB categories and fine-grained subtype discrimination. Because patients with overlapping liver disease etiologies were excluded, the reported performance may overestimate diagnostic accuracy in broader real-world clinical settings where overlapping syndromes are common. Multicenter external validation and prospective evaluation in more heterogeneous patient populations are needed before clinical implementation.

Read PDF

Similar papers

Sep 2026

Construction and evaluation of a machine-learning-based prediction model for pneumonia in patients with acute leukemia.

An interpretable XGBoost model that accurately predicts pneumonia risk in AL patients based on routine admission data is developed and validated and provides actionable risk stratification to inform preemptive diagnostic and therapeutic strategies.

W. Zhuang, Chen Huang, Xu-Dong Ma et al. · 0 citations
Review Open access Sep 2026

Steatosis Liver Index: A Validated Machine Learning Model Using Clinical Data Distinguishes Steatotic Liver Disease Phenotypes

Introduction and Objectives: Steatotic liver disease (SLD) is a leading cause of liver disease worldwide. Itsstratified into metabolic dysfunction-associated steatotic liver disease (MASLD), metabolic dysfunction and alcohol-associated liver disease (MetALD), and alcohol-associated liver disease (ALD) based on self-rep...

Mangesh Pagadala, Talal Khurshid, Wan-Yu Zhang et al. · 0 citations
Open access Sep 2026

Missingness-aware machine learning using routine laboratory data for distinguishing hepatitis from cirrhosis

Early identification of cirrhosis among patients with hepatitis is important for timely risk stratification and clinical management, yet conventional non-invasive scores such as the fibrosis-4 index (FIB-4) and the aspartate aminotransferase-to-platelet ratio index (APRI) may provide limited discrimination. In clin...

Qin Yang, Xue-Rui Hu, Shi-Hong Yu et al. · 0 citations
Open access Sep 2026

Leveraging time-series electronic health records with large language models for chronic kidney disease diagnosis in primary care

Abstract Objectives Chronic kidney disease (CKD) presents a growing public health challenge in China, exacerbated by low patient awareness and limited nephrology resources. This study evaluated the potential of large language models (LLMs) to support CKD diagnosis in primary care using time-series electronic health rec...

Jing-Yi Wu, Yue-Wen Zheng, Zhi-Jun He et al. · 0 citations
Open access Aug 2026

Prediction of Liver Disease Using Deep Learning and Artificial Neural Networks

Abstract— Liver disease is a major cause of illness and death worldwide, and early diagnosis remains a challenge due to the limited availability of trained medical professionals for interpreting diagnostic tests. This paper presents a deep-learning-based approach for the prediction of liver disease using the Indian Liv...

M.Dhana Laxmi, S. Nikitha · 0 citations
Open access Sep 2026

Machine learning models for predicting liver cancer: a real-world cohort study in China

The findings support the feasibility of leveraging large-scale inpatient laboratory data for risk-stratification model development and show good discrimination and interpretability for liver cancer prediction in a hospitalized real-world cohort using routinely available clinical and laboratory data.

Ce-Xiong Fu, Fang Li, Shi-Bing Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.