The findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings and should assess emerging high-parameter models and explore additional clinical domains.
Abstract
Extracting structured data from electronic health records (EHRs) remains a major challenge, particularly in non-English and resource-constrained healthcare systems. This study benchmarks multiple large language models (LLMs) for the automated extraction of structured clinical variables from Portuguese-language medical notes under limited computational resources. We evaluated five LLMs (GPT-4o mini, DeepSeek-V3, Mixtral-8x7B, LLaMA 8B, and Qwen-32B) against a manually curated dataset of cardiology and infectiology outpatient records. Models were deployed in quantized versions to optimize computational efficiency. Outputs were compared with human annotations using F1 score, balanced accuracy, and recall. Among the tested models, Qwen-32B achieved the highest performance in both the infectiology domain (balanced accuracy = 0.91 [0.07]) and cardiology domain (balanced accuracy = 0.89 [0.07]). Performance varied by clinical variable, with better results for frequently and consistently documented conditions (e.g., diabetes) and lower accuracy for complex or infrequent variables (e.g., tumors). Extraction time ranged from 0.9 to 24.2 minutes per patient, depending on clinical domain and model. These findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings. Future research should assess emerging high-parameter models and explore additional clinical domains.
Open LLMs can extract clinical findings from Finnish pediatric records with accuracy comparable to published English benchmarks, and uncertainty-based triage substantially reduces required expert workload.
J. Leinonen, J. Knuutila, S. Kurki et al.· medRxiv· 0 citations
BACKGROUND
Real-world data (RWD) in unstructured electronic health records (EHRs) is crucial for understanding complex diseases like cancer, but extracting structured information is challenging due to linguistic variability, semantic complexity, and privacy concerns. This study evaluates the performance of four small, locally deployable language models for information extraction from Italian EHRs.
METHODS
We examine three prompting strategies (zero-shot, few-shot, and annotated few-shot) across English and Italian, involving clinicians with varying expertise to assess the impact of prompt design on accuracy. We evaluate the performance of four open-source small language models (SLMs) for clinical information extraction from Italian electronic health records (EHRs) in the APOLLO 11 trial on non-small cell lung cancer (NSCLC). The extraction protocol involves four steps: problem definition, data preprocessing, Large Language Model (LLM)-based information extraction, and output evaluation.
RESULTS
We show that general-purpose models (e.g., LLaMA 3.1 8B) outperform biomedical models in most tasks, particularly in extracting binary features. Multiclass variables such as TNM (Tumor, Node, Metastasis) staging, PD-L1 (Programmed death-ligand 1), and ECOG-PS (Eastern Cooperative Oncology Group-Performance Status) are more difficult due to implicit language and lack of standardization. Few-shot prompting and native-language inputs significantly improve performance and reduced hallucinations. Clinical expertise enhances consistency in the extraction, particularly among students using annotated examples.
CONCLUSIONS
The study confirms that privacy-preserving SLMs can be deployed locally for efficient and secure cancer data extraction. Findings highlight the need for hybrid systems combining SLMs with expert input and underline the importance of aligning clinical documentation practices with SLM capabilities. This is the first study to benchmark SLMs on Italian EHRs and investigate the role of clinical expertise in prompt engineering, offering valuable insights for the future integration of SLMs into real-world clinical workflows.
Federica Corso, V. Peppoloni, L. Mazzeo et al.· Communications Medicine· 0 citations
Objectives
Social determinants of health (SDOH) are incompletely captured in structured electronic health records (EHRs) but are frequently documented in unstructured clinical notes. We evaluated large language models (LLMs) for extracting SDOH from clinical text.
Materials and Methods
We constructed an adult sepsis cohort from the Medical Information Mart for Intensive Care-IV (Sequential Organ Failure Assessment ≥2) and analyzed clinical notes from 1 year prior to 30 days following suspected infection. Three instruction‑tuned, decoder‑only LLMs (Mistral‑Instruct‑7B-v0.2, DeepSeek‑R1‑Distill‑Qwen‑14B, and GPT‑oss‑20B) were evaluated using structured prompts with predefined label schemas and few‑shot examples. Performance was benchmarked against a clinically validated annotated dataset and compared with a fine‑tuned encoder‑decoder baseline. Macro‑F1 scores were reported. A gold‑standard Intensive Care Unit (ICU) sepsis subset was independently annotated by 3 reviewers to assess domain‑level performance and ensemble strategies.
Results
Decoder‑only models outperformed the fine‑tuned encoder‑decoder baseline across SDOH domains. GPT‑oss achieved the highest macro‑F1 score (0.79) compared with Flan‑T5‑XXL (0.57). Prompt refinement substantially improved extraction accuracy. Ensemble majority voting increased robustness across domains, while unanimous agreement yielded high precision but limited coverage. In a subsequent mortality analysis, extracted SDOH did not independently predict 30-day mortality, which was instead associated with established clinical and demographic risk factors.
Discussion
Instruction‑tuned decoder‑only LLMs can reliably extract multiclass SDOH from unstructured clinical notes without task‑specific fine‑tuning. Ensemble and agreement‑based strategies provide practical operating points for high‑precision clinical deployment.
Conclusion
These findings support the feasibility of leveraging LLMs to enrich EHRs with structured SDOH data, providing a scalable approach for incorporating social context into downstream risk stratification and health outcome prediction.
D. Salazar, Pankaj Dipankar, Daniel Smolyak et al.· JAMIA Open· 0 citations
Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.
L. Barrett, N. Joshi, A. S. North et al.· medRxiv· 0 citations
PURPOSE
Manual chart review (MR) of electronic health records (EHRs) is time-consuming, error-prone, and limits the reproducibility and scalability of real-world data (RWD) research. Automation and standardization using natural language processing (NLP) could improve efficiency and scalability. CTcue is an NLP-based software platform designed to extract structured and unstructured data from EHRs. This study evaluated the accuracy and efficiency of CTcue versus MR in patients with early-stage resectable non-small cell lung cancer (NSCLC).
METHODS
Included were all patients with stage I to III NSCLC who underwent lung resections between January 2018 and December 2021 at the Leiden University Medical Center, the Netherlands. Demographics, tumor characteristics, treatment, and outcomes were collected. CTcue performance was compared with MR using weighted F1-scores, accuracy, precision, and recall for categorical variables and Bland-Altman analysis for continuous variables.
RESULTS
Eighty-five patients (70.2% of patients from the manual cohort) were identified by both methods and included in the comparison. CTcue achieved weighted F1-scores >0.85 for seven of 15 categorical variables, including sex, tumor location, and deceased status, although some scores were based on low number of observations in both cohorts. Lower performance was observed for variables with varying terminology in documentation, such as Eastern Cooperative Oncology Group status and pathological N-stage. Continuous variables showed negligible mean differences, indicating good agreement. Survival outcomes were identical in both data sets.
CONCLUSION
CTcue performance for patient selection was lower than anticipated. However, it enables accurate, efficient extraction of structured and unstructured EHR data in early-stage NSCLC. Manual validation remains necessary for variables with varying terminology. Further development of artificial intelligence-based tools-particularly for free-text data extraction-will be crucial to enhance the accuracy and scalability of future RWD research.
Hanieh Abedian Kalkhoran, Tobias Martinot, Lydia H N Schonewille et al.· JCO Clinical Cancer Informat...· 0 citations
This study provides a systematic comparison of several open-source LLMs on a structured intestinal symptom extraction task and concludes that Qwen3 models offer a favorable balance between accuracy and efficiency, making them suitable for resource-constrained scenarios.
Xinyue Zhang, Quanyu Wang, Beibei Liu et al.· Journal of Medical Internet...· 0 citations