Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world discharge summaries, and to identify recurring failure modes that limit reliability at scale. Materials and Methods: We applied a two-stage LLM pipeline---open-ended candidate identification (Gemini 2.5 Pro) followed by context-grounded verification (Gemini 2.5 Flash)---to 3,000 randomly sampled MIMIC-IV-Note discharge summaries. A subset of the pipeline output was then reviewed manually by clinical experts. Results: Our pipeline surfaced 3,460 candidate inconsistencies, affecting 69.7% of admissions. Representative examples spanned demographics, allergies, procedures, diagnoses, laboratory, medications, and care-planning domains, with direct implications for clinical reasoning or patient safety. Expert review also revealed recurring failure modes that arise when verification requires temporal reasoning, evolving-diagnosis context, or knowledge of outpatient-prescribing conventions the model does not natively possess. Discussion: Detection is highly context-dependent: many flagged pairs require anchoring each statement to its source section and clinical domain, then assessing whether the conflict reflects a true contradiction or missing context. We propose a graded ontology spanning strict contradiction and ambiguity, with a schema characterizing each flagged case by category, section, domain, and inconsistency axis. Conclusion: This formative study establishes a methodological foundation and conceptual framework to guide subsequent validated, large-scale EHR-inconsistency analysis.
Jiankang Lu, Panyu Chen, Miriam Treggiari et al.· 0 citations
Abstract Objectives We extracted a validated disease activity measure in rheumatoid arthritis (RA), the Clinical Disease Activity Index (CDAI), from a large tertiary academic medical center electronic health record (EHR) using an automated large language model (LLM)-based approach without requiring model pretraining. Materials and Methods The New York Presbyterian/Columbia University Medical Center Clinical Data Warehouse contains EHR data for over 4.5 million patients. RA patients were identified using International Classification of Disease-9 (ICD-9) and ICD-10 codes. Expert-curated CDAI keywords were extracted from unstructured notes using an automated natural language processing (NLP) pipeline leveraging GPT-4o API, a HIPAA-compliant, institutionally approved LLM platform. Performance was evaluated against expert chart review. Results Among 2756 RA patients with notes, 1038 (37.7%) were seropositive, 796 (28.9%) were seronegative, and 922 (33.4%) had unknown serostatus. Clinical Disease Activity Index and its components were extracted in 15.4% (160/1038) of seropositive patients indicating remission or low disease activity. Clinical Disease Activity Index documentation was more frequent among patients with multiple notes and among faculty, with high extraction accuracy (precision/recall/F1 = 0.97). Discussion This represents the first attempt to employ a zero-shot, ChatGPT-powered LLM platform to extract RA disease activity measures from real-world EHR data. Although a low prevalence of documentation was noted, important distinctions were observed when patients were subgrouped by serostatus, level of training, and number of visits. Conclusion An LLM-based pipeline accurately extracted CDAI from a single large academic EHR, revealing infrequent real-world documentation.
This entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.
Qiao Jin, Nicholas Wan, Robert Leaman et al.· Nature Protocols· 1 citation