Skip to content
Open access

Expert-Rated Documentary Quality of AI-Assisted Hospital Discharge Reports: A Retrospective Paired Comparison with Physician-Written Reports

Jul 2026 · Healthcare · Vol 14, pp. 2290 · 0 citations · 20 references
Medicine

TL;DR

AI-assisted discharge reports received higher expert-rated documentary quality scores in a non-blinded paired evaluation across most evaluated dimensions, supporting the need for a supervised hybrid model in which AI generates the initial draft while the clinician mandatorily validates sensitive content.

Abstract

Background/Objectives: The hospital discharge report is a critical document for care continuity that generates a substantial administrative burden for clinicians. Generative artificial intelligence (AI) offers the potential to reduce this burden while improving documentary quality. This study aims to compare, under real-world conditions with a GDPR-oriented architecture based on prior local anonymisation, the quality of AI-assisted discharge reports (IAIA) against those drafted by the responsible physician (INF). Methods: A retrospective, paired, expert-evaluation study was conducted at a Spanish university hospital. One hundred and twenty consecutive clinical cases from nine departments were included (240 reports total). Each case was independently evaluated by one of ten primary care physicians using a structured rubric covering 13 clinical dimensions (ordinal scale 1–3) and a global rating scale (1–10). The Wilcoxon signed-rank test was applied to all paired comparisons; effect size was estimated using the paired rank-biserial correlation (r). Results: IAIA achieved a significantly higher overall mean rating than INF (8.14 vs. 7.30 out of 10; p < 0.0001; r ≈ 0.76, large effect). IAIA was nominally superior in 9 of 13 clinical dimensions; after Bonferroni correction for the 13 per-dimension comparisons, six of these differences remained statistically significant, with the largest gains in family history, principal diagnosis hierarchy, and structured listing of secondary diagnoses. INF retained an advantage only in allergies and intolerances (2.69 vs. 2.46; p = 0.002), where IAIA tended to use generic formulas. Three dimensions showed no significant difference (prior treatment, physical examination, procedures). Conclusions: AI-assisted discharge reports received higher expert-rated documentary quality scores in a non-blinded paired evaluation across most evaluated dimensions. The physician-written report retained an advantage only in the safety-critical allergy domain, where allergy information must not be inferred by the model but sourced from verified structured fields or explicitly flagged as pending physician validation, supporting the need for a supervised hybrid model in which AI generates the initial draft while the clinician mandatorily validates sensitive content. Prior local anonymisation constitutes a GDPR-oriented approach to generative AI deployment in European hospital settings, substantially reducing the risk of disclosure of identifiable clinical information.

Read PDF

Similar papers

Review Open access Jul 2026

AI-Assisted Clinical Data Abstraction From Electronic Health Records: Retrospective Concordance Study

Abstract Background Manual chart abstraction from electronic health records is a critical step in clinical outcomes research but is time-intensive and prone to human error. Advances in artificial intelligence (AI), particularly large language models, offer the potential to automate the extraction of structured data from unstructured clinical documentation with improved efficiency and consistency. Objective This study aimed to evaluate the accuracy and efficiency of an AI-assisted approach for extracting patient-reported outcomes from clinical notes compared with traditional human abstraction. Methods We conducted a retrospective study of 26 patients treated with low-dose radiation therapy for osteoarthritis. Human reviewers abstracted numeric rating scale (NRS; 0‐10) pain scores at baseline, the end of treatment, and the first follow-up, and von Pannewitz score (VPS; 0‐4) improvement scores at posttreatment time points. A HIPAA (Health Insurance Portability and Accountability Act)–compliant generative pretrained transformer–based AI system was prompted to extract the same end points from clinical notes. Concordance was assessed using exact match rates, the intraclass correlation coefficient for the NRS, and weighted Cohen κ for the VPS. The time required for AI vs manual abstraction was recorded. The AI system was not trained or fine-tuned on study data, and performance was evaluated directly against human abstraction to reflect real-world deployment. Results The AI system demonstrated high concordance with human abstraction, achieving an exact match rate of 92% for the NRS (95% CI 84‐96; intraclass correlation coefficient=0.96) and 94% for the VPS (95% CI 84‐98; κ=0.91). All discrepancies were minor, and no spurious values were generated. The AI system identified 1 clinically relevant data point missed during manual review. Average abstraction time per patient decreased from approximately 30 minutes to 2 minutes, representing time savings of >90%. The system also captured trends in analgesic use, but these results were not statistically significant, including reductions without escalation. Conclusions AI-assisted data abstraction demonstrated high concordance with human review in this single-institution cohort while substantially reducing the time requirements. These findings support the feasibility of AI-assisted abstraction workflows, although further validation across larger and more diverse datasets is needed.

Camille Sarah Schwartz, M. J. Anderson, K. Moakler et al. · 0 citations
Review Open access Jul 2026

Artificial Intelligence-Generated Electronic Medical Record Summarization in Breast Surgical Oncology.

BACKGROUND Reviewing pathology, imaging, and consultation documents in oncology can be time-consuming, particularly when records originate from external facilities in different file formats. This study aimed to evaluate the impact of a Retrieval-Augmented Generation (RAG)-enabled GPT-4o summarization agent on clinical workflows and quality of outside-record summaries in breast surgical oncology. METHODS Initial performance evaluation of a GPT-4o/RAG agent to generate summaries of oncologic reports in 50 charts followed by a prospective pilot test of sequential cases, with each AI summary evaluated using a modified Provider Documentation Summarization Quality Instrument (PDSQI-9; 1-5 Likert scale), including dichotomized ratings (low [1-3], high [4, 5]), binomial testing, frequency and type of user-reported errors, clinician-coded error criticality (treatment-impacting vs noncritical). Pre- and post-use survey of documentation burden (NASA TLX) and user experience was performed. RESULTS Among 62 cases, AI-generated summaries were rated high for accuracy, usefulness, succinctness, and source citation. Thoroughness without omission was rated low in 28 (45%) summaries. Errors were noted in 25 (40%) surveys, with 13 (52%) classified as critical (treatment-impacting). The most common error type involved imaging, reported in 17 (68%) cases. For perceived time savings, the median response was neutral, but qualitative feedback described the tool as helpful for straightforward cases and as reducing typing burden but requiring workflow adjustment and improvements for complex cases. CONCLUSIONS Although users rated RAG-enabled GPT-4o agent-generated documentation summaries favorably on several quality domains, they frequently lacked thoroughness and occasionally contained treatment-relevant errors. Human review and further iteration of the technology remain necessary before implementation.

Ko Un Park, Bergen K. Sather, A. Shah et al. · 0 citations
Review Aug 2026

Clinical evaluation of a vision-language model for optimizing triage and clinical workflows in critical care.

OBJECTIVE Clinical deterioration in hospitalized patients is often preceded by subtle, dynamic physiological changes that are difficult to detect using intermittently charted electronic health record (EHR) data. Our objective was to evaluate the reliability, interpretability, and clinical relevance of a Vision‑Language Model (VLM)-based triage framework that analyzes physiological trend images, by comparing VLM-generated outputs with attending physician assessments as the expert clinical comparator. MATERIALS AND METHODS We conducted a single-center expert agreement pilot study including 100 adult patients with two hours of dynamic monitoring data across four vital signs (SpO2, RR, HR, BP). A structured prompt was developed using the Gemini 2.5 Flash model. Two independent reviewers assessed VLM outputs for clinical interpretation, artifact detection, and triage classification. We evaluated reviewer agreement using percent agreement and weighted Cohen's κ. A secondary risk-oriented analysis measured classification concordance and over-triaged cases as lower-risk classifications, while cases underestimating patient acuity were designated as higher-risk misclassifications. RESULTS The VLM demonstrated moderate to substantial agreement in triage classification with the attending physician and a low rate of under-triage. The VLM assigned the same patient acuity category as attending physician in 75 % of cases and underestimated acuity in 7 % of cases, compared with 14 % underestimation by the physician in training. DISCUSSION VLMs extend generative artificial intelligence capabilities by enabling image‑grounded clinical reasoning and offer signal‑processing capabilities for interpreting time‑stamped physiological trends. CONCLUSION The VLM demonstrated reliable clinical interpretation and an acceptable safety profile, however its integration into clinical workflows for early recognition of physiological deterioration and patient acuity assessment requires further rigorous evaluation and comparison to currently used track-and-trigger systems and patient monitoring methods.

I. Strechen, P. Krishnan, O. Kilickaya et al. · 0 citations
Review Open access Jul 2026

The High-Volume OPD Problem: Why Indian Clinical Documentation Requires a Purpose-Built Artificial Intelligence Model

Background: Ambient artificial intelligence (AI) clinical documentation platforms have demonstrated significant capacity to reduce physician documentation burden. However, existing commercial platforms are engineered for Western clinical ecosystems featuring 15-to-30-minute monolingual consultations. This narrative review evaluates the structural mismatches encountered when translating these architectures directly into the high-volume, short-duration, multilingual realities of Indian outpatient departments (OPDs). Methods: Electronic databases (PubMed, Scopus, IndMED, Google Scholar) were queried for clinical validation data, automatic speech recognition (ASR) performance metrics, and national healthcare workforce bulletins published between January 2017 and May 2026. A total of 23 core references were synthesized to map current systemic operational constraints. Results: The evaluation identified three acute structural mismatches: 1. A temporal conflict where natural language processing (NLP) architectures fail to reliably extract structured entities from compressed 1.9-to-6.9-minute Indian consultations. 2. A linguistic barrier where monolingual English ASR models suffer a 30% to 50% surge in word error rates when processing localized code-switched (Hinglish) dialogue. 3. An infrastructural disconnect due to the heterogeneous, often paper-based, electronic health record (EHR) footprint across Indian hospitals. Conclusion: Globally imported ambient documentation tools are structurally incompatible with Indian outpatient workflows. Resolving physician burnout securely requires establishing an India-native clinical AI research infrastructure optimized for short-consultation contexts and multi-language code-switched speech processing. Keywords: Ambient clinical intelligence, Clinical documentation, Outpatient department, Automatic speech recognition, Hinglish.

Anushtha Rakesh Chillure · 0 citations
Open access Mar 2026

Generating Patient Documents from Electronic Health Records Using Generative Artificial Intelligence: A Feasibility Study in a Japanese Cancer Center

Abstract Objectives Clinical documentation consumes substantial clinician time, potentially detracting from patient care. Generative artificial intelligence (AI) may support drafting discharge summaries and patient referral documents, but feasibility in non-Western-language oncology settings using real-world electronic health record (EHR) data remains insufficiently evaluated. This study assessed feasibility in a Japanese cancer hospital using an enterprise AI system. Methods Medical records from 61 consenting adult patients at Chiba Cancer Center were analyzed. Although the plan aimed at comprehensive EHR data, actual input was limited to extractable text (physician notes, nursing records); structured laboratory data and imaging, endoscopy, and pathology reports were not directly used, and existing summaries and external referrals were excluded to avoid information leakage. Data were converted to JavaScript Object Notation; GaiXer generated 31 discharge summaries and 30 referral documents. Four evaluators scored them; ≥80/100 was an exploratory threshold for draft-level practical utility. Feedback drove one refinement cycle. Results Generated documents scored approximately 60 to 70. A score ≥80 was reached by 9 of 31 discharge summaries in each evaluation; for referrals, none reached the threshold initially, whereas 5 of 30 did after refinement. Discharge summary scores did not substantially improve; referral scores did. Raw percent agreement among three nonphysician evaluators was high, although chance-corrected agreement varied. Wilcoxon signed-rank tests showed no significant change for discharge summaries ( p  = 0.866) but significant improvement for referrals ( p  = 0.006). Conclusion This feasibility study suggests AI may support drafting these documents in a secure environment using real-world Japanese EHR data, although the generated documents did not consistently reach the predefined threshold for draft-level utility. Findings should not be interpreted as demonstrating workload reduction or maximum performance under ideal data conditions. Future studies should evaluate larger datasets, multiple institutions and models, blinded evaluations, actual editing time, clinician acceptance, and workflow impact.

N. Michihata, Hiroshi Ishii, H. Tsujimura et al. · 0 citations
Open access Jul 2026

Effect of evaluation prompt strategies on LLM-as-a-judge reliability in critical care.

Bottom-up incremental scoring showed the closest alignment with human assessment in clinical AI evaluation, underscoring the need for standardised prompt architectures in clinical AI evaluation.

Jia-Yu Yan, Wing-Sum Chan, Ching-Tang Chiu et al. · 0 citations