In this simulation, AI-generated notes scored higher on documentation quality, varied less, and carried fewer clinically significant errors than notes written on the same consultations by junior-to-middle-grade clinicians.
Abstract
Background Ambient AI documentation tools, known as scribes, are entering routine clinical practice at scale, but the evidence comparing the notes they produce against clinician-written notes is dominated by single-site, single-language studies that rely on human review to find errors, a method known to miss most documentation errors. Methods We conducted a paired simulation across five countries and languages (Cambridge/English, Barcelona/Spanish, Milan/Italian, Paris/French, Cologne/German; 385 paired consultations, 770 notes). From each actor-performed consultation, an AI scribe (Heidi) and a junior-to-middle-grade clinician independently produced a note. Notes were scored on the PDQI-9 by evaluators blinded to authorship. Documentation errors were identified by two methods of deliberately different sensitivity - clinician adjudication, and a calibrated automated reviewer externally validated against a blinded ten-clinician panel - then graded for clinical risk by a three-model panel. The co-primary outcomes were PDQI-9 total and Critical+High error burden, the latter reported under both detection arms. The analysis plan was registered before any pooling across sites. Results AI notes scored higher than clinician notes on the PDQI-9 (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; Cohen dz=0.55), consistently across all five sites (dz 0.41-0.75), and were less dispersed (5.7% of AI vs 27.8% of clinician notes fell below the study pre-specified low-score threshold (<32)). On the principal safety outcome - the paired probability that a note carried [≥]Critical+High error - clinician notes were affected more often under both detection arms: 61.0% versus 24.4% by the calibrated reviewer (relative risk 2.50, 95% CI 2.09-3.00) and 21.8% versus 6.2% by clinician adjudication (relative risk 3.50, 95% CI 2.32-5.27). The difference was largest for omissions. Unaided clinician review identified roughly 12% of the errors the calibrated reviewer retained, and a smaller fraction in AI notes than in clinician notes. Conclusions In this simulation, AI-generated notes scored higher on documentation quality, varied less, and carried fewer clinically significant errors than notes written on the same consultations by junior-to-middle-grade clinicians. The magnitude of the safety difference depends on the sensitivity of error detection, so we report both detection regimes and bound rather than point-estimate the absolute error rate. Extension to live practice, consultant-authored documentation, and notes as filed after clinician editing remains to be established.
Objectives Safety claims for ambient artificial intelligence (AI) scribes rest on automated judges that detect documentation errors and grade clinical risk. Expert reviewers are under-sensitive and disagree with one another, so no gold standard exists and validation cannot mean accuracy. We tested whether such judges are a defensible instrument: reproducible, within the envelope of expert disagreement, and non-differential across arms. Methods Pre-registered, blinded validation study nested in a multi-country simulation of ambient AI documentation (English setting), reported per GRRAS. Ten external clinicians independently adjudicated a stratified sample of 434 pipeline flags, retained and screen-discarded, blinded to note authorship, identification source, the pipeline's verdict and severity tier. Agreement used Gwet's AC1; proportions carry Wilson intervals. Three propositions were pre-specified: envelope parity, non-differential behaviour across arms, and concordance on consensus cases. Results All ten reviewers completed: 565 adjudications across 434 items, 131 of them double-rated. Inter-clinician agreement on genuineness was fair (raw 59%, 95% CI 50 to 67; AC1 0.24), leaving no human consensus to serve as truth. Judge-clinician agreement was 64% (95% CI 60 to 68), overlapping that interval. Behaviour was near-symmetric on contrast-critical metrics: kept-precision 74% for AI against 81% for clinician notes, and severity signed gap +0.06 against -0.09 tiers. One sub-metric was asymmetric: removed-confirmed 56% against 42%, so the screen over-removes more on clinician notes, a direction conservative to the parent contrast. On 77 consensus items the pipeline concurred on 70% (95% CI 59 to 79). Latent-class triangulation placed the genuine-error rate among flagged candidates at 68% (94% credible interval 48 to 83). Conclusions The judges behave as a consistent, near-non-differential, clinician-equivalent instrument. This licenses a directional AI-versus-clinician contrast under a non-differential misclassification argument, subject to its conditions. It is not a claim of accuracy, which moderate consensus concordance and fair reliability preclude, and the genuine-error rate is best reported as an interval.
H. Bergman, V. Liu, B. Austin et al.· medRxiv· 1 citation· ⚡1
Background Prediction models are central to advancing precision oncology, yet many fail to translate into clinical practice due to methodological flaws and inadequate validation. This review provides a practical, clinician-oriented guide to the statistical principles and advanced methods for developing, validating, and interpreting robust prediction models. Methods This narrative review used a targeted literature search of PubMed, Embase, and Web of Science to identify methodological papers, reporting guidelines, and representative oncology prediction model studies, with a focus on literature published between January 1, 2005, and February 28, 2025. Landmark methodological papers published before 2005 were also included when directly relevant. Rather than performing a systematic review or meta-analysis, we synthesized key statistical principles and illustrative examples to guide clinicians and researchers through model development, validation, interpretation, and clinical translation. Findings A multifaceted evaluation encompassing discrimination, calibration, clinical utility, and external validation is essential for prediction models. Over-reliance on discrimination metrics such as the area under the receiver operating characteristic curve (AUC), while neglecting calibration and clinical utility, can lead to misleading conclusions about a model’s value. Rigorous external validation in geographically or temporally distinct cohorts is the most direct test of generalizability, and performance degradation should be interpreted through root-cause analysis rather than treated simply as model failure. Key challenges include managing overfitting, selecting appropriate modeling and validation strategies for different oncology scenarios, addressing special settings such as rare tumors and real-world data, and improving the interpretability of complex “black-box” models. Conclusion Building a trustworthy prediction model requires a combination of advanced computational methods and rigorous statistical principles. To bridge the gap from model development to clinical impact, researchers must prioritize comprehensive validation, transparent reporting, scenario-appropriate modeling decisions, and critical assessment of a model’s real-world utility.
Xuexing Wang, Youxian Dou, Yufeng Wang et al.· Frontiers in Oncology· 1 citation
These preliminary findings suggest that structured prompt engineering may meaningfully improve LLM-assisted diagnostic reasoning and warrant further investigation through prospective, multi-center trials with blinded adjudication.
Aya Kawssan, A. Bazzal, Aktham Abdelhadi et al.· Journal of Medical Research...· 0 citations
AI-assisted discharge reports received higher expert-rated documentary quality scores in a non-blinded paired evaluation across most evaluated dimensions, supporting the need for a supervised hybrid model in which AI generates the initial draft while the clinician mandatorily validates sensitive content.
Daniela Velásquez-Villegas, Toni Alonso Solís, Alex Trejo-Omeñaca et al.· Healthcare· 0 citations
BACKGROUND
The creation of diagnostic criteria for a disease is always challenging in medicine. Consensus meetings and expert's opinions are sometimes responsible for establishing useful and precise diagnostic criteria but, unfortunately, not all the consensus meetings are based on evidence-based conclusions.
METHODS
A perspective analysis of the historical evolution of diagnostic criteria for Vogt-Koyanagi-Harada.
RESULTS
We examine the historical evolution of diagnostic criteria for Vogt-Koyanagi-Harada disease, an autoimmune condition targeting melanocyte-containing tissues, primarily affecting the choroid and potentially involving the skin, ears, and meninges. Early diagnostic criteria proposed in the late twentieth century were considered inadequate, leading to an international consensus conference in 1999 and the publication of revised diagnostic criteria in 2001. However, these criteria contained major conceptual flaws, notably the combination of acute and chronic clinical features that rarely occur simultaneously. In addition, important diagnostic tools such as indocyanine green angiography were largely excluded due to prevailing local practices and skepticism regarding their use. These limitations hindered accurate diagnosis and delayed recognition of VKH as a disease with distinct acute-onset and chronic forms requiring different diagnostic frameworks and management strategies. Subsequent studies corrected some deficiencies but often generated complex criteria difficult to apply in routine practice.
CONCLUSION
We would like to emphasize the risks of poorly structured consensus processes and propose safeguards to ensure scientifically rigorous and clinically useful recommendations.
C. Herbort, Ioannis Papasavvas, A. A. El-Asrar et al.· Journal of Ophthalmic Inflam...· 0 citations
Bottom-up incremental scoring showed the closest alignment with human assessment in clinical AI evaluation, underscoring the need for standardised prompt architectures in clinical AI evaluation.