Skip to content
Review Open access

Augmenting structured diagnoses through effective use of pre-trained large language models on clinical notes

Sep 2026 · JAMIA Open · Vol 9 · 0 citations · 45 references
Medicine

TL;DR

LLMs represent a viable and flexible approach to diagnosis code extraction from unstructured clinical notes that can augment structured diagnoses and provide contextualizing metadata.

Abstract

Abstract Objective Clinical narrative provides a unique window into provider reasoning and attribution for automated diagnosis assignment, but large language models (LLMs) have traditionally not performed well at medical coding. We evaluate a reproducible method for automated diagnosis assignment using LLMs in clinical notes and compare with structured diagnoses. Materials and Methods We used GPT-OSS for prompt engineering and task segmentation to create a model that extracts ICD-10-CM diagnoses, with estimates of severity, currency, and importance, from progress notes. We assessed performance across multiple cohorts of patients aged 0-21 years. For each, 100 outpatient provider notes were selected across levels of severity, along with coded diagnoses from that visit (electronic health record [EHR]); a subset of 130 notes were subjected to clinical expert review. Results Comparison showed 18.7% exact code and 33.3% ICD-10-CM category match between EHR and LLM, but semantic similarity of 0.93 at the category level. Compared to expert review, LLM precision was 0.84 and recall 0.49 for exact matches, and 0.92 and 0.62, respectively, for category-level matching. In contrast, coded diagnoses showed slightly higher precision (0.94 for both cases) and substantially lower recall (0.27 and 0.43) versus expert review. Codes not identified by the LLM were more often rated by the reviewer as lower importance or certainty. Discussion We demonstrate a reusable approach to optimize LLMs for use in diagnosis extraction from clinical notes that can augment structured diagnoses and provide contextualizing metadata. Conclusion LLMs represent a viable and flexible approach to diagnosis code extraction from unstructured clinical notes.

Read PDF

Similar papers

#natural language process... Preprint Aug 2026

KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs

This work proposes Knowledge-Guided Reasoning over Clinical Evidence with LLMs (KREL), a framework that leverages LLMs for clinical text understanding and reasoning while integrating external ICD coding guidelines as structured knowledge, and enables tight coupling between domain knowledge and LLM reasoning.

Xubin Chen, Yi-Peng Zhou, Wenxin Sun et al. · 0 citations
Open access Sep 2026

Clinical Code Mapping with LLM Tool Use: A Pilot for Automated Data Extraction of Medication and Diagnosis Information from Unstructured Clinical Notes.

LLMs are suitable for information extraction of medications from clinical notes for use in research databases, however, for a clinical setting where the treatment of patients would be dependent on LLM performance, the current state-of-the-art open weight models are not accurate enough.

T. Spreuer, A. Günther, R. Majeed · 0 citations
Open access Aug 2026

Evaluation of Diagnostic Accuracy of Open-Source and Proprietary Large Language Models Across Multi-System Clinical Cases

A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.

Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al. · 0 citations
Open access Aug 2026

Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals

Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.

L. Barrett, N. Joshi, A. S. North et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.