Skip to content
Review Open access

Large language model–based patient handoff tools relevant to intensive care: a scoping review

Aug 2026 · Considerations in Medicine · Vol 4, pp. e000046 · 0 citations · 13 references

TL;DR

Current evidence suggests that LLMs have considerable potential to support future ICU handoff workflows but substantial evidence gaps remain and prospective evaluation in real-world ICU settings, with clinically meaningful safety metrics and structured communication frameworks, is needed.

Abstract

Provider-to-provider patient handoffs are a routine yet complex component of intensive care unit (ICU) workflows and are essential for patient safety. Emerging generative artificial intelligence (AI), particularly large language models (LLMs), may improve clinical communication through automated summarisation, documentation and handoff-related workflows. This scoping review followed PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and searched six databases (PubMed, IEEE Xplore, Google Scholar, ACM Digital Library, arXiv and medRxiv) on 7 October 2025 for studies published between 14 August 1999 and 6 June 2025. Eligible publications examined LLMs or related language technologies in clinical summarisation, documentation, communication workflows, clinical decision support or patient handoffs. Forty-six studies met the inclusion criteria. Most evaluated LLMs for clinical summarisation, information extraction or decision support, whereas studies directly evaluating LLM-generated patient handoffs were rare. Only one study examined emergency department handoffs and no studies evaluated ICU-specific handoffs. Across the literature, LLMs demonstrated promising performance for clinical summarisation but recurrent challenges included hallucinations, clinically relevant omissions, limited model transparency and reliance on linguistic evaluation metrics rather than patient-centred outcomes. Human expert review was frequently incorporated to assess clinical accuracy and safety. Overall, current evidence suggests that LLMs have considerable potential to support future ICU handoff workflows but substantial evidence gaps remain. Prospective evaluation in real-world ICU settings, with clinically meaningful safety metrics and structured communication frameworks, is needed before widespread implementation.

Read PDF

Similar papers

Review Aug 2026

How large language models can be used for teamwork and communication in healthcare settings: A scoping review.

LLMs hold substantial potential to enhance healthcare teamwork by supporting clinical decisions, streamlining administrative workflows, and improving patient communication, however, ethical, legal, and accountability concerns remain.

Ilse Super, Olya Rezaeian, Onur Asan · 0 citations
Sep 2026

Clinical Practice Guidelines as a JSON Service: A Proposed Architecture for Decision Support in Major Burning Management

Clinical Practice Guidelines (CPGs) encode evidence-based clinical knowledge but are primarily distributed as unstructured PDF documents, making them inaccessible to automated clinical decision support (CDS) systems. This paper proposes a service-oriented architecture that formalizes the IMSS Clinical Practice Guideline for Major Burn Management (IMSS-375) as a versioned REST/JSON microservice. Seven clinical decision endpoints are defined, each encapsulating a specific GPC recommendation: burn classification, initial assessment, fluid resuscitation, pain management, infection prevention, nutritional support, and transfer criteria, following HL7 FHIR R4 interoperability standards. A mapping between GPC clinical rules (including the Parkland formula, Benaim scale, Curreri formula, and Baux prognostic index) and structured JSON request/response schemas is presented and evaluated against related formalization approaches and verified through structured schema-level invocations against a representative clinical scenario, including a detailed comparison of Mexico IMSS-375 standard properties against HL7 FHIR and OpenEHR. A deployment architecture is described covering hospital-level integration, a centralized service layer, and a non-relational persistence tier based on document-oriented storage for unstructured clinical data.

José de Jesús Álvarez Ramírez, Rocío Maciel, V. Larios · 0 citations
Review Open access Jul 2026

Transformer-Based Language Models for Clinical Decision Support Using Clinical Notes: A Scoping Review

Transformer-based language models have been applied across diverse clinical-note tasks, but the evidence base more strongly supports retrospective task feasibility than transportability, equitable performance, workflow benefit, or safe clinical deployment.

Saahoon Hong, Hunhui Na · 0 citations
Review Open access Aug 2026

Examining Electronic Medical Record Migration in a Team-Based Primary Care Clinic: Case Study

Abstract Background Electronic medical record (EMR) systems are now ubiquitous in Ontario primary care, with near-universal adoption among family physicians. EMRs have become critical pieces of clinic infrastructure that can affect nearly every aspect of clinical and administrative work. As EMR vendors consolidate and functionality evolves, some clinics undertake EMR migrations—complex transitions that require data conversion, workflow redesign, training, and change management. Despite their importance and wide adoption, there is a gap in understanding how community-based primary care organizations in Canadian settings navigate migration in practice, particularly from the perspective of frontline clinicians and staff. Objective The project aimed to describe a single case of EMR migration within an 11-physician, team-based Family Health Organization (FHO) in Peterborough, Ontario, from the perspective of frontline clinicians and staff. This included assessing the factors influencing the decision to migrate, the operational and human factors that shaped implementation, and practical lessons for supporting and evaluating future EMR transitions. Methods A single-site, physician-led quality improvement project was conducted in collaboration with the Peterborough Ontario Health Team using an exploratory case study approach. This approach was used to examine a complex, real-world health IT transition in its local organizational context. This project examined one FHO’s migration between EMR providers using purposive sampling of individuals directly involved in the transition, including a lead physician (n=1), a lead administrator (n=1), family physicians (n=9), interprofessional health care providers (n=12), and medical office assistants (n=9). Data sources included 2 semistructured interviews, 2 focus groups, and 1 comprehensive document review. Data collection spanned the planning, implementation, and stabilization phases of the migration from February 2022 to November 2024. Sessions were recorded, transcribed, and deidentified. Documents included emails, meeting minutes, training communications, project planning materials, and troubleshooting records. Interview and focus group transcripts were analyzed thematically. Study documents were reviewed separately to construct a narrative chronology of the migration. Results Four interrelated themes emerged: (1) burden of invisible labor, (2) cost and uncertainty as a catalyst, (3) importance of physician leadership, and (4) engage stakeholders early. Participants described migration as a resource-intensive organizational transition shaped by unpaid preparation work, cost pressures, clinical governance, quality assurance, and the responsiveness of internal and external stakeholders, particularly the vendor. Conclusions This study illuminates the process of an EMR migration within one team-based FHO. Effective transitions depend on resourcing the often-invisible work of change, sustaining physician-led governance, and securing early, collaborative engagement with vendors and external partners. These findings complement provincial migration guidance by emphasizing protected time for training and quality assurance, local workflow adaptation, staged validation to shorten the transition period (“gray zone”), and explicit assessment of vendor responsiveness. Lessons are readily actionable for similar group practices and clinics planning future migrations or optimizations.

Max Moloney, M. Sundareswaran · 0 citations
Review Open access Aug 2026

An Electronic Health Record–Integrated, Large Language Model–Powered Tool to Triage Surgical Patients

Key Points Question Can surgical patient triage be automated using a large language model (LLM) agentic workflow? Findings In this quality improvement study, the LLM tool recommended hospitalist consultation for nearly a quarter of the 6193 triaged cases. The tool achieved 94% sensitivity and 74% specificity, and post hoc medical record review suggested that most discrepancies reflected modifiable gaps in clinical criteria, institutional workflow, or physician practice variability, rather than LLM misclassification. Meaning The findings of this study suggest that an LLM-powered human-in-the-loop agentic workflow could accurately triage surgical patients for a surgical comanagement service.

Janelle B. Wang, T. Keyes, April S. Liang et al. · 0 citations
Review Open access Aug 2026

Evidence, use cases, and implementation safeguards of large language models in primary care

Recent developments in large language models (LLMs) have created new opportunities to support primary care, where much of clinical work is text-mediated. This narrative review synthesizes evidence on LLM applications relevant to primary care workflows and summarizes implementation safeguards. Across studies, the most consistently supported near-term value is workflow augmentation, particularly documentation and inbox management (e.g., drafting portal replies and summarizing information for clinician review) and communication support, where benefits are reported primarily as process endpoints (time, acceptability, perceived communication quality) rather than hard patient outcomes. Evidence for improvements in clinician diagnostic reasoning, treatment planning, and downstream patient outcomes is more limited and context-dependent, and many evaluations remain simulated or conducted in adjacent settings, limiting generalizability to routine primary care. Accordingly, potential roles in population health and cost reduction should be treated as hypothesis-generating and evaluated prospectively. Challenges related to privacy, security, transparency, and model reliability shape organizational governance requirements and evolving regulatory expectations for the clinical use of generative AI in primary care. We emphasize a pragmatic adoption approach: prioritize high-volume, lower-risk clerical and communication workflows; maintain clinician verification and accountability; and apply governance and equity safeguards (e.g. privacy, security, transparency, auditability, monitoring for drift and error) before scale-up. Christof et al. provide a narrative review that synthesizes evidence on LLM applications relevant to primary care workflows and summarizes implementation safeguards. They highlight the remaining need to demonstrate improved patient outcome of clinical improvement in many studies and outline a pragmatic adoption approach in practice.

Michael Christof, Krish Patel, Jiandong Zhou et al. · 0 citations