Skip to content
#generative ai Review Open access

Generative artificial intelligence in clinical reasoning and differential diagnosis in internal medicine.

Aug 2026 · Medicina clínica (Ed. impresa) · Vol 166 10, pp. 107567 · 0 citations · 44 references
Medicine

TL;DR

A narrative review of the available evidence presents a narrative review of the available evidence on the effect of LLMs on diagnostic reasoning, the optimal design of clinician-LLM interaction, the appropriate timing of consultation during the clinical encounter, the safest models of clinical-AI integration, and the main risks associated with their use.

Abstract

Clinical reasoning and differential diagnosis are core competencies in medicine. Large language models (LLMs) have generated considerable interest as potential tools to support these skills. This article presents a narrative review of the available evidence, organized around five key questions: the effect of LLMs on diagnostic reasoning, the optimal design of clinician-LLM interaction, the appropriate timing of consultation during the clinical encounter, the safest models of clinical-AI integration, and the main risks associated with their use. The evidence shows that LLMs improve differential diagnosis when used by trained professionals within structured workflows. However, passive use generates biases, and clinician-AI collaboration may not consistently outperform autonomous LLMs. A practical framework stratified by degree of diagnostic uncertainty is proposed, with operational and educational recommendations oriented toward "physician-in-the-loop" models, in which LLMs amplify, challenge, and make explicit the diagnostic reasoning process under critical human oversight.

Read PDF

Similar papers

Review Open access Aug 2026

Large Language Models for Differential Diagnosis: A Survey of Performance, Collaboration, and Technical Strategies

Errors in differential diagnosis often arise while clinicians are generating and comparing candidate explanations. This review examines the use of large language models (LLMs) for this part of diagnostic reasoning. Internal medicine and pediatrics are the main focus; evidence from radiology, surgical subspecialties, infectious disease, and mental health is used to examine how findings change across specialties. Reported performance depends on the clinical setting, the quality of the input, the prompt, model adaptation, and the evaluation design. Some studies place LLMs near trainees and find that they produce wider, better-organized differentials. Experienced clinicians, however, remain more reliable overall. Domain adaptation, external knowledge, and interactive workflows have improved performance in specific evaluations, but hallucinations and automation bias remain, alongside unresolved questions of governance. Current evidence therefore supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.

Yun-Jia Wu, Qi Yan, Dingcheng Tian · 0 citations
Review Open access Aug 2026

Large Language Models and Medical AI Systems for Healthcare Diagnosis: A Systematic Review

Although promising, LLM-based systems are not yet reliable enough for autonomous medical diagnosis, and multiple recommendations for future research are contained to ensure a high level of safety, transparency, and clinical applicability for LLMs and other AI/ML-related technologies and devices.

M. U. K. Gunawardhna, Pirunthavi Wijikumar, D. Weerasinghe · 0 citations
Review Open access Jul 2026

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

A dual-view approach that connects clinical practice with computational methods is presented, establishing a five-level competency scheme following Miller’s Pyramid and linking deductive, inductive, and abductive reasoning patterns to common medical goals and tasks.

Qi Peng, Jiatong Li, Sirui Huang et al. · 5 citations
Review Open access Aug 2026

Generative artificial intelligence in medical education: from knowledge assessment to clinical reasoning and professional competence

Generative artificial intelligence (GenAI), particularly large language models (LLMs), is poised to fundamentally transform medical education. Based on a structured literature search of PubMed, Scopus, Web of Science, and Google Scholar, this review synthesizes current evidence on the applications, capabilities, and limitations of GenAI across the medical training continuum. Advanced LLMs demonstrate a formidable command of medical knowledge, consistently achieving passing scores on standardized licensing examinations, with GPT-4 and domain-specific models like Ortho GPT showing particular proficiency. As versatile teaching tools, these models can generate high-quality assessment materials, provide personalized on-demand tutoring, and power interactive virtual patients for clinical reasoning practice. However, this potential is tempered by significant challenges, including a propensity for “hallucinations,” embedded biases that can perpetuate health inequities, linguistic performance disparities, and a fundamental gap in flexible, adaptive clinical reasoning. Integration also raises critical concerns regarding academic integrity, potential over-reliance leading to deskilling, and data privacy. Responsible adoption requires a structured approach encompassing the development of tiered AI competency frameworks, blended curricular integration, dedicated faculty development, and a rigorous research agenda focused on longitudinal learning outcomes. Ultimately, GenAI should be viewed as a powerful augmentative tool, not a replacement for human educators. Its successful integration will depend on leveraging its strengths to enhance efficiency and scalability while preserving the essential humanistic elements of medical practice through expert oversight and validation.

R. Xie, Bei-En Zhang, Lifeng Xiao · 0 citations
Preprint Jul 2026

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop, and the evidence of safety does not yet exist.

Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem et al. · 0 citations
Review Jul 2026

Can AI assist in reducing diagnostic error? A narrative review

It is concluded that AI tools have matured to the extent that they can improve diagnostic decision-making of clinicians and can assist institutions in increasing diagnostic safety.

Ian A. Scott · 0 citations

Related blog posts