Skip to content
Open access

The potential of LLMs in generating questions and answers with EHRs

Jul 2026 · Frontiers in Digital Health · Vol 8 · 1 citation · 32 references
Medicine

TL;DR

Conventional medical education requires clinicians to formulate questions and answers based on prototypes from EHRs, which is heuristic and time-consuming, this study shows that mainstream LLMs could generate questions and answers with real-world EHRs at levels close to clinicians.

Abstract

Background This study aimed to generate medical qualification exam questions and their corresponding answers from real-world electronic health records (EHRs) with large language models (LLMs), and to compare their output to that of human medical experts. Methods Utilizing a multicenter bidirectional anonymized database China Elderly Comorbidity Medical Database (CECMed), a total of 8 LLMs: ERNIE 4, ChatGLM 4, Doubao, Hunyuan, Spark 4, Qwen, Llama 3, and Mistral were tasked with generating open-ended questions and answers based on a subset of sampled admission reports. LLMs generated the medical question and answer through few-shot prompting. An independent expert panel scored the AI-generated outputs based on multiple criteria, including coherence, sufficiency of key information, information correctness, factual consistency, evidence of statement, and professionalism, using 5-point Likert scales. Results For question generation, ERNIE 4 achieved the highest cumulative score (16.47). Human experts surpassed LLMs in sufficiency of key information (3.67) but lagged in information correctness (3.63 vs. LLMs' 4.03–4.57). The information correctness of ERNIE was significantly higher than the human's [0.93 (0.62, 1.24), p < 0.01]. For answer generation, humans led overall (14.49), while Doubao outperformed the other LLMs in coherence (3.57), factual consistency (3.60), and professionalism (3.53). The coherence of human's was significantly better than that of 8 LLMs, especially outperformed Llama [0.8 (0.37, 1.23), p < 0.01] and Mistral [0.87 (0.45, 1.28), p < 0.01]. Conclusions Conventional medical education requires clinicians to formulate questions and answers based on prototypes from EHRs, which is heuristic and time-consuming. This study shows that mainstream LLMs could generate questions and answers with real-world EHRs at levels close to clinicians. Although current LLMs performed dissatisfactorily in some aspects, medical students and interns may find LLMs a useful auxiliary tool to support their learning. Clinical Trial Registration https://clinicaltrials.gov/study/NCT06316544, identifier: NCT06316544.

Read PDF

Similar papers

Open access Jul 2026

Large language models for interpretation of health checkup results

Overall, advanced prompting markedly improved model performance, and top-tier LLMs demonstrated robust interpretive capability, while caution is needed for variables with complex clinical semantics such as blood pressure.

Jiwon You, Hangsik Shin · 0 citations
Preprint Aug 2026

KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs

This work proposes Knowledge-Guided Reasoning over Clinical Evidence with LLMs (KREL), a framework that leverages LLMs for clinical text understanding and reasoning while integrating external ICD coding guidelines as structured knowledge, and enables tight coupling between domain knowledge and LLM reasoning.

Xubin Chen, Yipeng Zhou, Wenxin Sun et al. · 0 citations
Aug 2026

Large Language Models as Clinical Support Tools in Drug Information Services: Performance Comparison With Pharmacists.

LLMs show potential as tools for preliminary drug information retrieval and rapid responses generation in drug information services, however, variable concordance and persistent limitations in citation credibility indicate the need for continued pharmacist oversight.

Nuntapong Boonrit, Najwa Bin-Useng, Aphichaya Sirijariyawat et al. · 0 citations
Open access Jul 2026

Clinician expertise and prompt engineering enhance cancer information extraction in electronic health records by small language models.

BACKGROUND Real-world data (RWD) in unstructured electronic health records (EHRs) is crucial for understanding complex diseases like cancer, but extracting structured information is challenging due to linguistic variability, semantic complexity, and privacy concerns. This study evaluates the performance of four small, locally deployable language models for information extraction from Italian EHRs. METHODS We examine three prompting strategies (zero-shot, few-shot, and annotated few-shot) across English and Italian, involving clinicians with varying expertise to assess the impact of prompt design on accuracy. We evaluate the performance of four open-source small language models (SLMs) for clinical information extraction from Italian electronic health records (EHRs) in the APOLLO 11 trial on non-small cell lung cancer (NSCLC). The extraction protocol involves four steps: problem definition, data preprocessing, Large Language Model (LLM)-based information extraction, and output evaluation. RESULTS We show that general-purpose models (e.g., LLaMA 3.1 8B) outperform biomedical models in most tasks, particularly in extracting binary features. Multiclass variables such as TNM (Tumor, Node, Metastasis) staging, PD-L1 (Programmed death-ligand 1), and ECOG-PS (Eastern Cooperative Oncology Group-Performance Status) are more difficult due to implicit language and lack of standardization. Few-shot prompting and native-language inputs significantly improve performance and reduced hallucinations. Clinical expertise enhances consistency in the extraction, particularly among students using annotated examples. CONCLUSIONS The study confirms that privacy-preserving SLMs can be deployed locally for efficient and secure cancer data extraction. Findings highlight the need for hybrid systems combining SLMs with expert input and underline the importance of aligning clinical documentation practices with SLM capabilities. This is the first study to benchmark SLMs on Italian EHRs and investigate the role of clinical expertise in prompt engineering, offering valuable insights for the future integration of SLMs into real-world clinical workflows.

Federica Corso, V. Peppoloni, L. Mazzeo et al. · 0 citations
Preprint Aug 2026

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower, and this results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.

Praveen Reddy, C. Mandke, Suvrankar Datta et al. · 0 citations