The authors' analysis reveals that LLMs demonstrate promising capabilities in processing textual and visual data related to various liver diseases, including hepatocellular carcinoma, cirrhosis, and non-alcoholic fatty liver disease, but study heterogeneity and significant challenges remain regarding accuracy, reliability, and safety.
Abstract
Background and Aim The rapid advancement of generative artificial intelligence (AI), particularly large language models (LLMs), has opened new frontiers in healthcare, with emerging implications for hepatology. This systematic review synthesizes the current state of research on the application of LLMs in hepatology, focusing on their capabilities in real-world clinical settings, limitations, and future directions. Materials and Methods Electronic databases, including MEDLINE, EMBASE, and OVID as a search platform, were used to identify eligible studies from inception to January 2025. Eligible studies investigated the clinical utility and performance of LLMs in hepatology, with a clear comparison to a defined ground truth. Key findings were extracted and synthesized narratively. The ROBINS-I tool was used to assess the risk of bias in each study. Results Twenty-one studies were included in this review. Our analysis reveals that LLMs demonstrate promising capabilities in processing textual and visual data related to various liver diseases, including hepatocellular carcinoma, cirrhosis, and non-alcoholic fatty liver disease. LLMs effectively assisted with radiological image interpretation, provided clinical decision support, and generated patient education materials. However, the accuracy of these models was highly variable, depending on the specific task and the complexity of the clinical scenario. Limitations, such as the generation of inaccurate or misleading information (“hallucinations”), dependence on training data quality, and ethical considerations, were identified across multiple studies. Conclusion Generative AI demonstrates feasibility across various hepatology applications, but study heterogeneity and significant challenges remain regarding accuracy, reliability, and safety. Future integration necessitates further research into training methods, data quality, ethical considerations, and real-world validation against standardized benchmarks.
Current evidence indicates that LLMs have substantial potential to enhance healthcare delivery, research, and personalized medicine, but they should currently be regarded as supportive tools rather than autonomous clinical decision-makers.
Antoni Klamka, Paulina Kawalec, Kamil Bronikowski et al.· 0 citations
A scoping review of 24 PubMed-indexed studies published between 2023 and 2026 was conducted to assess current applications, benefits, limitations, and future directions of LLMs in healthcare.
Antoni Klamka, Paulina Kawalec, Kamil Bronikowski et al.· Quality in Sport· 0 citations
Artificial intelligence (AI), particularly foundation and generative models, is reshaping the practice of hepatology through enhanced knowledge synthesis, quantitative and reproducible analysis of multimodal data, and personalized clinical decision support. This narrative review examines the transition from task-specific discrimination AI to large language models (LLMs), multimodal foundation models, and agentic AI. We synthesize evidence from original and validation studies, clinical evaluations, and benchmark studies, as well as expert reviews and regulatory frameworks across metabolic dysfunction-associated steatotic liver disease, chronic hepatitis B, cirrhosis and portal hypertension, hepatocellular carcinoma, and liver transplantation. LLMs can convert free-text notes into structured data, summarize longitudinal electronic health records, support patient education, and retrieve guideline-based information. Retrieval-augmented generation and agentic AI may improve traceability and workflow support, but current evidence is largely retrospective or proof-of-concept. In digital pathology and imaging, discriminative AI has enabled more quantitative and reproducible histologic scoring and biomarker analysis. Pathology and multimodal foundation models offer transferable representations, report generation, and cross-modal reasoning, but hepatology-specific validation remains limited. Key risks include hallucination, automation bias, domain shift across centers and devices, and inequities due to under-representation of patient subgroups. We outline the future directions for safe AI model deployment based on multimodal foundation models, prospective and federated evaluation, lifecycle governance, and continuous monitoring for performance, calibration, and equity. Most generative AI applications in hepatology remain at the proof-of-concept stage, and rigorous prospective validation with human-in-the-loop oversight is required before clinical integration.
Nana Peng, Mary Yue Wang, S. J. Song et al.· Clinical and Molecular Hepat...· 0 citations
Aims: To systematically review large language model (LLM) and natural language processing (NLP) studies published in first-quartile (Q1) clinical radiology journals, focusing on methodological quality, model implementation, and comparative performance. Methods: A systematic search of PubMed and Scopus was conducted to identify original studies involving LLMs or transformer-based NLP systems published in Q1 clinical radiology journals through June 20, 2025. Eligible studies were screened and assessed for methodological characteristics, including dataset type, involving imaging modality (if any), model used, model accessibility, prompt disclosure, and handling of stochasticity. Human-LLM/NLP and LLM/NLP-LLM/NLP performance comparisons were extracted. Results: Fifty-six studies were included, most published in 2024-2025. Proprietary models such as GPT-4 and GPT-4o were most frequently evaluated. Real-world clinical data were used in 62.5% of studies, but only 10.7% reported a power analysis, and 39.1% addressed stochasticity. Prompt engineering was reported in 41.9% of studies. In 455 human-LLM/NLP comparisons, LLMs/NLPs outperformed humans in 54 cases, while humans outperformed in 79; most results (70.8%) were ties. Among 3,164 valid LLM/NLP-LLM/NLP comparisons, GPT-4o had better performance than earlier models. Conclusion: LLMs/NLPs demonstrated performance comparable to radiologists in many text-based tasks but remain inconsistently evaluated. Methodological limitations, including lack of power analysis, incomplete reporting, and under-addressed stochasticity, hinder robust assessment. Greater transparency, standardized evaluation protocols, and inclusion of diverse clinical settings are essential for reliable integration into radiology practice.
I. Mese· Journal of Health Sciences a...· 0 citations
It was revealed that most GenAI applications in healthcare rely on general purpose LLMs to provide treatment recommendations, and future research should prioritize the development of interpretable, domain-specific models and rigorous clinical trials to ensure safe and effective integration into healthcare settings.
Leonides Medeiros Neto, Maicon Herverton Lino Ferreira da Silva Barros, Kayo Henrique de Carvalho Monteiro et al.· Journal of Healthcare Inform...· 0 citations
A conceptual Clinical Co-pilot Framework is proposed to position GenAI as a collaborative partner that supports clinicians rather than replaces them, which provides a conceptual basis for future empirical validation and may help inform the responsible implementation of GenAI in healthcare.
Lina Cheng, Chia-Yu Hung, Te-Nien Chien· International Journal of Adv...· 0 citations