Sep 2026· Journal of Managed Care & Specialty Pharmacy· Vol 32 9, pp.
1062-1075
· 0 citations· 41 references
Medicine
TL;DR
LLMs hold considerable promise for automating methodological appraisal of RWE studies; however, their performance is variable and model dependent.
Abstract
Background
Real-world evidence (RWE) is increasingly used to inform regulatory and payer policy decisions and health technology assessment, yet appraising the methodological credibility of RWE studies remains time-intensive and requires specialized expertise. The appraisal task could involve using an appraisal tool that provides a structured approach for evaluating bias in observational studies of comparative effectiveness and safety. Large language models (LLMs) may offer a scalable means to support this appraisal process, but their performance on structured bias assessment tasks has not been fully characterized.
Objective
To compare the performance of LLMs from 6 major artificial intelligence (AI) technology providers against human expert assessments in appraising bias in published RWE studies using the Appraisal of Potential Bias in Real-World Evidence Studies framework.
Methods
We conducted a comparative diagnostic accuracy study evaluating 40 LLMs from OpenAI, Anthropic, Google, xAI, Meta, and DeepSeek. Ten published RWE studies representing diverse pharmacoepidemiological designs and data sources were appraised by each LLM using a structured chain-of-thought prompt with conditional rubric injection based on the Appraisal of Potential Bias in Real-World Evidence Studies framework. Two independent human reviewers with pharmacoepidemiology training evaluated each study, with a third adjudicator resolving disagreements to establish the reference standard. LLM performance was assessed using overall accuracy and macro-averaged precision, recall, and F1 scores. Assessment time was compared between models and benchmarked against human reviewers. Bootstrap method was used to construct 95% CI for performance measures.
Results
Across 280 item-level assessments per model (10 studies × 28 items), overall accuracy ranged from 12.9% to 66.1%. The highest-performing model was Claude-Sonnet-4.6 (66.1%), followed by o3 (65.4%) and Gemini-3.1-pro-preview (65.0%). Macro-averaged F1 scores ranged from 30.9% to 66.9%; o3 achieved the highest F1 score (66.9%), followed by GROK-4 (65.5%) and Gemini-3.1-pro-preview (65.4%). Human reviewers required an average of 61.05 minutes per study; all LLMs completed assessments substantially faster, with average time per study ranging from 0.80 to 17.22 minutes relative to humans.
Conclusions
LLMs hold considerable promise for automating methodological appraisal of RWE studies; however, their performance is variable and model dependent. Their greatest value may lie in enhancing efficiency and supporting human-led appraisal as decision-support tools rather than replacing expert review. Future research should assess performance across larger, more diverse RWE study collections and evaluate output reproducibility across repeated runs.
While unsuitable to be used as a sole assessor, ChatGPT-o3 may serve as an adjunct tool to enhance the efficiency and consistency of RoB assessments in systematic reviews.
Siddharth Gandhi, A. Shokravi, Y. Chelliahpillai et al.· PLoS ONE· 0 citations
Abstract Background A growing body of literature leverages large language models (LLMs) to make mental health predictions. However, these models are prone to bias, and studies to validate their clinical utility are lacking. Objective This scoping review aims to uncover bias and clinical utility limitations stemming from the methodological design of LLM-based mental health predictive systems. In addition, it intends to document the level of self-reflection about bias and clinical challenges reported by authors in their own work. Methods This work follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and was registered online. Eligible studies were original research articles in English published between 2019 and 2024, using LLMs to detect mental health conditions in nonsynthetic textual data. The search was conducted in 5 scientific databases (PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and ACL Anthology) with queries associating keywords related to “Mental Health,” “Large Language Models,” and “Prediction.” We extracted both methodological information about the included studies and authors’ statements relevant to issues of bias and clinical utility. This extraction was based on a framework screening the entire pipeline of development of LLMs with applications in mental health: research design and selection, data collection, outcome definition, model development, and postdeployment considerations. Statistical description of the retrieved entities, as well as thematic coding, was performed for analysis. Results A total of 2472 articles were identified, of which 263 (10.6%) were assessed for eligibility, and 201 (8.1%) were included in the review. Included studies were mostly recent, indicating a growing interest in the use of LLMs for mental health predictions. Our analysis revealed that a majority of studies share similar methodological choices along their development pipeline: most of them focus on depressive disorders identified via processing user texts on social media, mainly with the use of nonspecialist LLMs derived from BERT (Bidirectional Encoder Representations from Transformers). Following previous works on these matters, we highlighted how these choices may hinder the clinical relevance and fairness of the envisioned systems. Similarly, we found that 164 (81.6%) studies mention themes related to bias and clinical utility; however, most of the discussion revolves around data-centered issues. Only 41 (20.4%) articles mention themes associated with at least 3 out of 5 pipeline steps, suggesting a limited appropriation of the notions of bias and clinical utility in such a sensitive context as mental health analysis. Conclusions Bias and clinical utility are lightly covered in the field of LLM-based mental health prediction research as of 2019‐2024. In-depth approaches involving interdisciplinary teams of clinicians and natural language processing specialists are needed to ensure technical soundness, clinical relevance, and fair outcomes for potential users.
Clémentine Bleuze, Karen Fort, Vincent P. Martin et al.· JMIR AI· 0 citations
It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.
Euijun Yang, S. Ko, Hyekyung Woo· Journal of Medical Internet...· 0 citations
OBJECTIVE
Data extraction is among the most resource-intensive and error-prone stages of systematic review production. Large language models (LLMs) offer potential for automating or semi-automating this process, yet their performance characteristics remain incompletely characterised. This systematic review aimed to comprehensively evaluate LLM accuracy, reliability, and efficiency for data extraction in evidence synthesis, and to identify optimal implementation strategies.
METHODS
We searched PubMed, Embase, Web of Science, and preprint servers (medRxiv, arXiv) through December 2025. Studies were eligible if they evaluated one or more LLMs for data extraction against a human reference standard and reported quantitative performance metrics. Two reviewers independently extracted data and assessed methodological quality using PROBAST + AI and reporting completeness using TRIPOD-LLM. Narrative synthesis was performed due to substantial heterogeneity precluding meta-analysis.
RESULTS
Twenty-seven studies met inclusion criteria, evaluating models including GPT-4/4o (n = 15), Claude versions 2-3.5 (n = 10), Gemini (n = 3), and open-source alternatives including Llama, Mistral, Qwen, and DeepSeek. Overall accuracy ranged from 47% to 99.9%, with substantial heterogeneity by task type and data granularity. Categorical and string variables were extracted more reliably (74-96%) than numerical data (47-88%). Claude 3.5 Sonnet achieved high accuracy in an assistive workflow (91.0%; 95% CI: 90.4-91.6%), exceeding human-only extraction (89.0%). Claude models outperformed GPT in head-to-head comparisons (OR 1.70 for event counts). Omissions were the dominant error type (60-74%), with hallucination rates of only 0.08-6%, challenging widespread fabrication concerns. Time savings of 33% to 87% were reported, although most included studies did not quantitatively assess efficiency. Methodological quality was generally robust, with 74.1% of studies rated low risk of bias under PROBAST + AI. Mean TRIPOD-LLM compliance was 88.5%, though gaps in inference settings and model version documentation were common.
CONCLUSION
LLMs demonstrate promising but variable performance for data extraction in evidence synthesis. Current evidence supports their integration as assistive tools within dual-extraction workflows requiring human verification, rather than as autonomous extractors. Categorical data is extracted more reliably than numerical outcomes, and few-shot prompting with structured output formats consistently improves performance. Standardised benchmarks and prospective comparative studies remain priorities for future research.
Ravi Shankar, Amaevia Lim, Xu Qian· Journal of Biomedical Inform...· 0 citations
Background While health technology assessment (HTA) acceptance of contextual real-world data (RWD) studies describing burden of disease, disease natural history, or treatment pathways is relatively common, HTA practices for RWD studies addressing real-world clinical efficacy, such as those using external control arms (ECAs), are still evolving and less standardized. The aim of this study was to use data from HTA submissions and reports to understand common analytical methods and data considerations for submissions using RWD-based ECA. This evaluation used ECA studies as a basis for investigating the use of RWD to evaluate clinical efficacy. Methods Secondary data were compiled from selected oncology submissions to HTA agencies between January 2016 and December 2022 that incorporated RWD-based ECA data, using natural language-processing text-mining to identify and select relevant cases. Submissions were reviewed in six countries across Asia Pacific (Australia), Europe (France, Germany, UK), and North America (Canada and US). Submissions that were rated both positive and negative by HTA agencies were included, with HTA feedback organized into generalizability, confounding, data quality, and data analysis categories. Results Of 204 submissions identified, 100 cases were selected for the analysis of patterns highlighting sources of data for ECAs and RWD methodology best practices: Australia (n = 3), Canada (n = 34), France (n = 19), Germany (n = 15), UK (n = 26), and US (n = 3). A positive HTA recommendation was received by 69 of these 100 cases. Lung cancer was associated with the greatest number of cases/submissions. Retrospective cohort studies were the most common source of RWD, with inverse probability of treatment weighting/propensity score weight as the most common methodology used to generate real-world evidence. Most of the selected RWD-based ECA cases were from Canadian and UK HTA agencies. Positive comments focused on population adjustment, RWD viability, and alignment of data with standard of care (SoC) for that country/indication; negative comments focused on missing/limited data, lack of alignment with SoC, and potential risk of bias. Conclusion This study captured challenges in considering RWD-based ECAs for HTA submission and presents criteria for creating viable RWD studies using ECAs. Data source selection, patient population comparison, and transparent presentation of potential biases were important factors in enhancing the credibility and utility of RWD-based ECAs in HTA decision-making processes.
Gleicy Macedo Hair, M. Hanisch, S. Mt-Isa et al.· Oncology Reviews· 0 citations