Skip to content
Review

Uncertainty-aware extraction of clinical findings from Finnish EHRs using open large language models

Jul 2026 · medRxiv · 0 citations
Medicine

TL;DR

Open LLMs can extract clinical findings from Finnish pediatric records with accuracy comparable to published English benchmarks, and uncertainty-based triage substantially reduces required expert workload.

View source

Similar papers

Open access Jul 2026

Benchmarking large language models for clinical data extraction from Portuguese medical notes in a university hospital

The findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings and should assess emerging high-parameter models and explore additional clinical domains.

Breno Gabriel Araújo Sampaio de Jesus, Tomaz Castrillon Figueiredo, Clariele de Almeida Pereira et al. · 1 citation
Open access Jul 2026

Clinician expertise and prompt engineering enhance cancer information extraction in electronic health records by small language models.

BACKGROUND Real-world data (RWD) in unstructured electronic health records (EHRs) is crucial for understanding complex diseases like cancer, but extracting structured information is challenging due to linguistic variability, semantic complexity, and privacy concerns. This study evaluates the performance of four small, locally deployable language models for information extraction from Italian EHRs. METHODS We examine three prompting strategies (zero-shot, few-shot, and annotated few-shot) across English and Italian, involving clinicians with varying expertise to assess the impact of prompt design on accuracy. We evaluate the performance of four open-source small language models (SLMs) for clinical information extraction from Italian electronic health records (EHRs) in the APOLLO 11 trial on non-small cell lung cancer (NSCLC). The extraction protocol involves four steps: problem definition, data preprocessing, Large Language Model (LLM)-based information extraction, and output evaluation. RESULTS We show that general-purpose models (e.g., LLaMA 3.1 8B) outperform biomedical models in most tasks, particularly in extracting binary features. Multiclass variables such as TNM (Tumor, Node, Metastasis) staging, PD-L1 (Programmed death-ligand 1), and ECOG-PS (Eastern Cooperative Oncology Group-Performance Status) are more difficult due to implicit language and lack of standardization. Few-shot prompting and native-language inputs significantly improve performance and reduced hallucinations. Clinical expertise enhances consistency in the extraction, particularly among students using annotated examples. CONCLUSIONS The study confirms that privacy-preserving SLMs can be deployed locally for efficient and secure cancer data extraction. Findings highlight the need for hybrid systems combining SLMs with expert input and underline the importance of aligning clinical documentation practices with SLM capabilities. This is the first study to benchmark SLMs on Italian EHRs and investigate the role of clinical expertise in prompt engineering, offering valuable insights for the future integration of SLMs into real-world clinical workflows.

Federica Corso, V. Peppoloni, L. Mazzeo et al. · 0 citations
Open access Aug 2026

Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals

Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.

L. Barrett, N. Joshi, A. S. North et al. · 0 citations
Review Open access Jul 2026

AI-Assisted Clinical Data Abstraction From Electronic Health Records: Retrospective Concordance Study

Abstract Background Manual chart abstraction from electronic health records is a critical step in clinical outcomes research but is time-intensive and prone to human error. Advances in artificial intelligence (AI), particularly large language models, offer the potential to automate the extraction of structured data from unstructured clinical documentation with improved efficiency and consistency. Objective This study aimed to evaluate the accuracy and efficiency of an AI-assisted approach for extracting patient-reported outcomes from clinical notes compared with traditional human abstraction. Methods We conducted a retrospective study of 26 patients treated with low-dose radiation therapy for osteoarthritis. Human reviewers abstracted numeric rating scale (NRS; 0‐10) pain scores at baseline, the end of treatment, and the first follow-up, and von Pannewitz score (VPS; 0‐4) improvement scores at posttreatment time points. A HIPAA (Health Insurance Portability and Accountability Act)–compliant generative pretrained transformer–based AI system was prompted to extract the same end points from clinical notes. Concordance was assessed using exact match rates, the intraclass correlation coefficient for the NRS, and weighted Cohen κ for the VPS. The time required for AI vs manual abstraction was recorded. The AI system was not trained or fine-tuned on study data, and performance was evaluated directly against human abstraction to reflect real-world deployment. Results The AI system demonstrated high concordance with human abstraction, achieving an exact match rate of 92% for the NRS (95% CI 84‐96; intraclass correlation coefficient=0.96) and 94% for the VPS (95% CI 84‐98; κ=0.91). All discrepancies were minor, and no spurious values were generated. The AI system identified 1 clinically relevant data point missed during manual review. Average abstraction time per patient decreased from approximately 30 minutes to 2 minutes, representing time savings of >90%. The system also captured trends in analgesic use, but these results were not statistically significant, including reductions without escalation. Conclusions AI-assisted data abstraction demonstrated high concordance with human review in this single-institution cohort while substantially reducing the time requirements. These findings support the feasibility of AI-assisted abstraction workflows, although further validation across larger and more diverse datasets is needed.

Camille Sarah Schwartz, M. J. Anderson, K. Moakler et al. · 0 citations
Review Open access Jul 2026

A Large Language Model Leaderboard for Clinical Note Entity Extraction

An LLM leaderboard showing how open-source LLMs perform at entity extraction on unseen clinical notes is developed, showing that large language models are already available that can perform entity extraction well enough to be considered in place of some administrative data.

E. Martin, Seungwon Lee, K. Riazi et al. · 0 citations
Open access Aug 2026

Evaluation of Diagnostic Accuracy of Open-Source and Proprietary Large Language Models Across Multi-System Clinical Cases

A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.

Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al. · 0 citations