Aug 2026· Journal of Healthcare Informatics Research· 0 citations· 46 references
TL;DR
It was revealed that most GenAI applications in healthcare rely on general purpose LLMs to provide treatment recommendations, and future research should prioritize the development of interpretable, domain-specific models and rigorous clinical trials to ensure safe and effective integration into healthcare settings.
Abstract
The rapid evolution of Generative Artificial Intelligence (GenAI) presents significant opportunities to transform healthcare, particularly in generating personalized treatment recommendations. This systematic literature review explores the current state of GenAI language models applications in various medical domains, assessing their effectiveness, applicability, and limitations. The review addresses nine specific research questions to understand the potential and challenges of integrating GenAI into clinical practice. We use the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. From a pool of 3237 studies, 42 were selected based on inclusion and exclusion criteria. These studies were analyzed to evaluate the use of generative language models, such as GPT-3 and GPT-4, in various medical domains including oncology, cardiovascular, gastrointestinal, and ophthalmological care. The analysis revealed that most GenAI applications in healthcare rely on general purpose LLMs to provide treatment recommendations. Fine-tuning with domain-specific data and prompt engineering were found to significantly improve output quality and reliability. However, persistent challenges include lack of clinical validation, ethical concerns such as bias, and issues related to transparency and regulatory compliance. While GenAI demonstrates strong potential to support clinical decision-making, real-world deployment remains limited due to unresolved ethical and validation issues. Future research should prioritize the development of interpretable, domain-specific models and rigorous clinical trials to ensure safe and effective integration into healthcare settings.
Current evidence indicates that LLMs have substantial potential to enhance healthcare delivery, research, and personalized medicine, but they should currently be regarded as supportive tools rather than autonomous clinical decision-makers.
Antoni Klamka, Paulina Kawalec, Kamil Bronikowski et al.· 0 citations
Although promising, LLM-based systems are not yet reliable enough for autonomous medical diagnosis, and multiple recommendations for future research are contained to ensure a high level of safety, transparency, and clinical applicability for LLMs and other AI/ML-related technologies and devices.
M. U. K. Gunawardhna, Pirunthavi Wijikumar, D. Weerasinghe· Sri Lankan Journal of Applie...· 0 citations
The authors' analysis reveals that LLMs demonstrate promising capabilities in processing textual and visual data related to various liver diseases, including hepatocellular carcinoma, cirrhosis, and non-alcoholic fatty liver disease, but study heterogeneity and significant challenges remain regarding accuracy, reliability, and safety.
T. Suenghataiphorn, Narisara Tribuddharat, Pojsakorn Danpanichkul et al.· Hepatology Forum· 0 citations
OBJECTIVES
Systematic literature reviews (SLRs) underpin life sciences research but are resource intensive. Generative artificial intelligence, particularly large language models (LLMs), may accelerate key SLR tasks, yet performance and reliability for evidence synthesis remain unclear. This manuscript aims to review current evidence on GenAI performance across core SLR tasks.
METHODS
We conducted a PRISMA-adapted rapid evidence assessment of English-language biomedical studies published from November 2022 to July 2025 evaluating GenAI or LLMs for systematic literature review tasks, including search strategy development, title/abstract screening, full-text screening, data extraction, risk-of-bias assessment, qualitative synthesis, report writing, and end-to-end review generation. Findings were summarized qualitatively by task.
RESULTS
Among 115 included studies, evidence supporting the use of GenAI was strongest for title/abstract screening (n=51) and data extraction (n=33). Selected high-quality evaluations reported sensitivities ≥90%, workload reductions of 27-71%, and human-comparable or superior performance in calibrated human-in-the-loop workflows. Evidence for full-text screening (n=15) and risk-of-bias assessment (n=17) was more variable, showing gains in structured or fine-tuned implementations but persistent limitations in specificity and nuanced judgment. For search strategy development, qualitative synthesis, and report writing, GenAI was most effective as a supportive tool; fully autonomous end-to-end SLR generation was unreliable.
CONCLUSIONS
GenAI can improve efficiency across multiple SLR tasks when used in hybrid human-AI workflows. Current evidence supports targeted, task-specific adoption with transparent reporting and human oversight, rather than full automation.
R. Fleurence, Riaz Qureshi, Rakesh Aggarwal et al.· Value in Health· 1 citation
A conceptual Clinical Co-pilot Framework is proposed to position GenAI as a collaborative partner that supports clinicians rather than replaces them, which provides a conceptual basis for future empirical validation and may help inform the responsible implementation of GenAI in healthcare.
Lina Cheng, Chia-Yu Hung, Te-Nien Chien· International Journal of Adv...· 0 citations
A scoping review of 24 PubMed-indexed studies published between 2023 and 2026 was conducted to assess current applications, benefits, limitations, and future directions of LLMs in healthcare.
Antoni Klamka, Paulina Kawalec, Kamil Bronikowski et al.· Quality in Sport· 0 citations