Skip to content
Open access

SynTrustBench: An Evidence-Gated and Executable Benchmark for Trustworthiness Claims in Synthetic Clinical Data

Aug 2026 · medRxiv · 0 citations
Medicine

TL;DR

SynTrustBench is introduced, an evidence-gated and executable benchmark for evaluating trustworthiness claims across five non-compensable dimensions: fidelity, clinical utility/validity, privacy, equity, and robustness/generalization.

Abstract

Synthetic clinical data are increasingly used for healthcare machine-learning development, model validation, data sharing, and predeployment testing, yet such data often claim to be trustworthy after passing a limited collection of realism tests. A synthetic dataset may indeed claim statistical similarity while leaking training membership, erasing rare subgroups, failing on held-out real patients, or lacking sufficient artifacts for reproduction. We introduce SynTrustBench, an evidence-gated and executable benchmark for evaluating trustworthiness claims across five non-compensable dimensions: fidelity, clinical utility/validity, privacy, equity, and robustness/generalization. Its Evidence Assessment component audits published reports and produces a five-element Evidence Maturity Profile (EMP) together with a separate evaluability gate. Its executable structured-tabular protocol accepts frozen real training data, held-out real test data, a synthetic table, and a declarative configuration; computes dimension-specific metrics and uncertainty; and produces subgroup results, failure flags, benchmark cards, and provenance manifests. In a frozen pilot audit of 30 reports, 17 of 30 quantitatively evaluated privacy, 2 of 30 documented a formal privacy guarantee to the audit threshold, 2 of 30 evaluated equity, 12 of 30 evaluated robustness, and only 4 of 30 passed the evaluability gate. The executable implementation operationalizes the same dimensions through distribution and dependency checks, frozen train-on-real/test-on-real (TRTR) and train-on-synthetic/test-on-real (TSTR) utility, empirical privacy attacks, subgroup analysis, perturbation testing, and a controlled failure-injection harness. SynTrustBench does not certify clinical safety or collapse trustworthiness into a single score. Instead, it provides an inspectable predeployment contract for identifying what was evaluated, what failed, what remains unknown, and whether evidence is sufficiently complete and reproducible for comparison or downstream healthcare AI use.

Read PDF

Similar papers

#machine learning Review Aug 2026

CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility

CoMedBench is introduced, a reproducible benchmark that evaluates a family of generators under a common clinical-validity framework and one shared training and evaluation engine, spanning static tabular and temporal downstream tasks on established critical-care datasets.

Akanta Das, Farhad Al-Amin Dipto, Mrinmoy Sarkar Anto et al. · 0 citations
Book Open access Aug 2026

LiveMedBench: A Contamination-Limited Medical Benchmark for LLMs with Automated Rubric Evaluation

The deployment of Large Language Models (LLMs) in high-stakes clinical settings demands rigorous and reliable evaluation. However, existing medical benchmarks remain static, suffering from two critical limitations: (1) data contamination, where test sets inadvertently leak into training corpora, leading to inflated performance estimates; and (2) temporal misalignment, failing to capture the rapid evolution of medical knowledge. Furthermore, current evaluation metrics for open-ended clinical reasoning often rely on either shallow lexical overlap (e.g., ROUGE) or subjective LLM-as-a-Judge scoring, both inadequate for verifying clinical correctness. % To bridge these gaps, we introduce LiveMedBench, a continuously updated, contamination-limited, and rubric-based benchmark that weekly harvests real-world clinical cases from online medical communities, ensuring strict temporal separation from model training data. We propose a Multi-Agent Clinical Curation Framework that filters raw data noise and validates clinical integrity against evidence-based medical principles. For evaluation, we develop an Automated Rubric-based Evaluation Framework that decomposes physician responses into granular, case-specific criteria, achieving substantially stronger alignment with expert physicians than LLM-as-a-Judge. % To date, LiveMedBench comprises 2,756 real-world cases spanning 38 medical specialties and two languages, paired with 16,702 unique evaluation criteria. Extensive evaluation of 38 LLMs reveals that even the best-performing model achieves only 39.2%, and 84% of models exhibit performance degradation on post-cutoff cases, confirming pervasive data contamination risks. Error analysis further identifies contextual application---not factual knowledge---as the dominant bottleneck, with 35-48% of failures stemming from the inability to tailor medical knowledge to patient-specific constraints. The code and data are available at https://github.com/ZhilingYan/LiveMedBench/ LiveMedBench.

Zhiling Yan, D. Song, Zhe Fang et al. · 0 citations
Open access Jul 2026

Addressing benchmarking gaps in large language models for health and medicine with dynamic red-teaming

A Dynamic, Automatic and Systematic red-teaming audit framework that continuously stress-tests LLMs for health across four safety-critical axes: robustness, privacy, bias and hallucination, which provides a scalable framework for surfacing latent risks before such systems are deployed in consumer-facing health assistants and broader clinical workflows.

Jiazhen Pan, Bailiang Jian, Paul Hager et al. · 0 citations
Review Aug 2026

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably. We introduce CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR), a benchmark for retrospective clinical audit: 25 clinician-validated scenarios instantiated as 750 patient-specific cases over real-patient-derived MIMIC-IV data. Systems investigate each case through a governed, logged tool environment for record retrieval, computation, and policy access, and return one of four verdicts---Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous---the last two separating missing evidence from residual medical ambiguity. Beyond verdict accuracy, we score patient-evidence and policy grounding, process adherence, calibrated abstention, reliability, and efficiency against case-level reference verdicts produced by independent multi-model adjudication and calibrated against Clinical Board review. Every retrieval, computation, and report is replayable, so the investigation trace is inspectable and scorable. To our knowledge, CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework. Across 16 agentic systems, four-way accuracy spans 65.3-76.1%, but raw accuracy overstates investigation quality. Defect-free accuracy, which credits a verdict only when correct and free of prohibited shortcuts, is 4.8-14.8 points lower and reorders the leaderboard.

Veronica Chatrath, Bryan Zhu, George Pu et al. · 0 citations
Preprint Jul 2026

ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science

CLINLENS is introduced, a benchmark of 200 executable tasks over five linked MIMIC resources spanning structured electronic health records, notes, electrocardiograms, chest radiographs, and echocardiograms, which exposes a substantial gap between runnable submissions and correct clinical analyses.

Yuan Zhu, Ethan B. Liu, Frank Nie et al. · 0 citations
#artificial intelligence Preprint Aug 2026

CareGraph: An Auditable Hybrid AI Framework for Evidence-Grounded Personalized Longitudinal Health Intelligence

CareGraph is an auditable hybrid AI framework that converts heterogeneous records into prioritized trends, missing context indicators, bounded next steps, discussion questions, and provenance linked explanations that offers a safety bounded foundation for intelligent personalized health systems.

Prof. S. B. Ghawate, Tanvi R. Patil · 0 citations