It is concluded that AI tools have matured to the extent that they can improve diagnostic decision-making of clinicians and can assist institutions in increasing diagnostic safety.
Abstract
Abstract Diagnostic error, defined as missed, wrong, or delayed diagnoses or those not communicated to patients, is common, affecting 5–10 % of hospital admissions and clinic visits. Such errors cause patient harm in up to 1 in 100 of such encounters and account for 10 % of all hospital deaths and serious adverse events. About 80 % of diagnostic errors are potentially preventable, most resulting from flaws in clinician reasoning in formulating and testing diagnostic hypotheses. The advent of artificial intelligence (AI), and large language models (LLMs) in particular, has attracted great interest in how these technologies can reduce diagnostic error within the context of bedside or clinic consultations. This narrative review aims to provide practising clinicians with a comprehensible analysis of where AI and LLMs are currently positioned in assisting diagnostic performance in clinician-patient encounters based on contemporary state-of-the-art research. It attempts to answer seven questions relevant to clinician understanding and adoption of AI/LLMs. It concludes that AI tools have matured to the extent that they can improve diagnostic decision-making of clinicians and can assist institutions in increasing diagnostic safety. The rapid development of LLMs and ongoing release of new versions necessitate continuous monitoring of their evolving diagnostic capabilities. Importantly, a balanced approach is required where LLMs work to complement, rather than replace, the nuanced diagnostic reasoning of clinicians.
A narrative review of the available evidence presents a narrative review of the available evidence on the effect of LLMs on diagnostic reasoning, the optimal design of clinician-LLM interaction, the appropriate timing of consultation during the clinical encounter, the safest models of clinical-AI integration, and the main risks associated with their use.
L. Corral-Gudino, M. Ramos-Casals, M. Marcos et al.· Medicina clínica (Ed. impres...· 0 citations
Abstract Objectives This study examines AI’s capacity to mitigate noise-related diagnostic errors, evaluates its impact on accuracy, and explores the interplay between AI-driven efficiency and human clinical reasoning, particularly in rare or complex cases. Background: Diagnostic errors in clinical reasoning are significantly influenced by noise – random unwanted variability in expert judgments – distinct from cognitive biases. Despite debiasing efforts, noise persists, contributing to adverse events. Artificial intelligence (AI) offers potential solutions but faces limitations in addressing novel diagnostic scenarios requiring creative reasoning. Methods A narrative review and conceptual analysis synthesises the literature on noise in medical decision-making, AI applications in healthcare, and clinical reasoning frameworks. Case studies (e.g., radiology, pathology) and empirical data on AI performance are reviewed, alongside discussions of noise types (occasion, pattern, group-level) and the role of AI in decision hygiene. Results AI reduces noise by minimising unwanted variability in pattern recognition tasks (e.g., imaging analysis), improving diagnostic consistency. However, AI struggles with novel hypotheses, creative reasoning, and contextual interpretation, remaining reliant on human oversight. AI’s inability to replicate human creativity limits its utility in rare disease diagnosis. Hybrid human-AI systems show promise but require balancing AI noise reduction capacities with human contextual judgment. Conclusions AI mitigates noise-driven errors in structured tasks but cannot replace human reasoning in complex, uncertain scenarios. Optimal diagnostic accuracy demands the integration of AI’s analytical strengths with clinicians’ creative and contextual reasoning. Future research should prioritise AI collaboration frameworks and address AI limitations in novelty-driven disease diagnostics.
Errors in differential diagnosis often arise while clinicians are generating and comparing candidate explanations. This review examines the use of large language models (LLMs) for this part of diagnostic reasoning. Internal medicine and pediatrics are the main focus; evidence from radiology, surgical subspecialties, infectious disease, and mental health is used to examine how findings change across specialties. Reported performance depends on the clinical setting, the quality of the input, the prompt, model adaptation, and the evaluation design. Some studies place LLMs near trainees and find that they produce wider, better-organized differentials. Experienced clinicians, however, remain more reliable overall. Domain adaptation, external knowledge, and interactive workflows have improved performance in specific evaluations, but hallucinations and automation bias remain, alongside unresolved questions of governance. Current evidence therefore supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.
Yun-Jia Wu, Qi Yan, Dingcheng Tian· AI Medicine· 0 citations
It is concluded that LLM-based decision-support tools hold substantial promise as complementary — rather than autonomous — decision-support systems capable of transforming medication safety and pharmacy practice.
K. K. Kumar, Koyya Gowtham Reddy, K. Reddy· International Scientific Jou...· 0 citations
This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop, and the evidence of safety does not yet exist.
Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem et al.· 0 citations
Abstract Artificial intelligence (AI) is entering clinical practice faster than the evidence base supporting it. Clinicians, who remain the licensed decision-makers at the bedside, increasingly find themselves as end-users of tools whose strengths, failure modes, and external validity they have had no opportunity to assess. This article offers a practical framework that does not require coding or mathematical literacy. We outline how AI is built, validated, deployed, and monitored, and where each phase typically goes wrong. We propose seven questions that clinicians can run through to evaluate any clinical AI tool in the time it takes to read an abstract, alongside a traffic-light schema for matching oversight to risk and a short list of demands clinicians should make of vendors and institutions. We then examine the deeper questions of equity, accountability, and the therapeutic relationship that AI is now forcing into view. AI literacy belongs alongside biostatistics and evidence-based medicine as a core clinical competency.
Alaa Abdelqader, M. Alkhateeb, Abdullah Al-Marrawi et al.· Avicenna Journal of Medicine· 0 citations