While unsuitable to be used as a sole assessor, ChatGPT-o3 may serve as an adjunct tool to enhance the efficiency and consistency of RoB assessments in systematic reviews.
Abstract
Objective To evaluate the reliability and diagnostic performance of ChatGPT-o3 in conducting Risk of Bias (RoB) assessments of randomized clinical trials (RCTs) using the Cochrane RoB 2.0 tool. Materials and methods This methodological validation study analyzed 50 RCTs sampled from 50 published meta-analyses. Each trial was independently assessed by the original systematic review authors (OSRAs), our masked human panel, and ChatGPT-o3. Structured prompts based on RoB 2.0 guidelines were used to elicit ChatGPT-o3 assessments. Agreement was evaluated using weighted Cohen’s kappa and Gwet’s AC2. Diagnostic performance was measured by sensitivity, specificity, and balanced accuracy, with human ratings as the reference. Results ChatGPT-o3 classified 34% of trials as high risk, compared with 22% by our panel, and 12% by the OSRAs. Agreement was modest (median κ: 0.33 with our panel; 0.14 with OSRAs). Overall Gwet’s AC2 was 0.30. For detecting high-risk trials, ChatGPT-o3 achieved a sensitivity of 0.46, specificity of 0.69, and balanced accuracy of 0.57. For low-risk trials, its sensitivity was 0.47, specificity was 0.86, and balanced accuracy was 0.66. Discussion The results indicate that ChatGPT-o3 produced more conservative RoB ratings than human reviewers, identifying a greater percentage of trials as having a high RoB. Conclusion While unsuitable to be used as a sole assessor, ChatGPT-o3 may serve as an adjunct tool to enhance the efficiency and consistency of RoB assessments in systematic reviews.
OBJECTIVE
The recently introduced ROBUST-RCT tool aims to reconcile ease of application with methodological rigor in risk-of-bias assessments for systematic reviews. The tool is structured in two steps: first, evaluating what happened; second, judging the risk of bias related to the assessed aspect of the study. Its straightforward design aims to avoid overly complex workflows. Thus, its usability testing included junior reviewers to ensure accessibility and ease of use. No data regarding its inter-rater reliability are currently available. This study aims to assess the inter-rater reliability of the ratings between junior researchers using the ROBUST-RCT.
STUDY DESIGN
An inter-rater reliability study. Four junior researchers screened and rated a random sample of 115 articles from a systematic search on PubMed. An additional phase beyond the prospectively defined research phases was introduced to exclude articles in which the two raters assessed different outcomes, resulting in a sample of 85 articles. As pre-specified in the protocol, the primary statistical analysis employed Gwet's AC2 at each step for each core item ("step-level") and at aggregated step 2 ratings ("judgment set") to provide an overview of the final judgment in the tool. Exploratory analyses include Fleiss' Kappa and a block-level approach.
SETTING
Universidade Federal do Rio Grande do Sul, a university in southern Brazil.
RESULTS
In the primary analysis, the aggregated data with the step 2 ratings ("judgment set") yielded a Gwet's AC2 agreement coefficient of 0.59 (95% CI: 0.53, 0.65); its inter-rater reliability was classified in Gwet's benchmarking as "moderate or higher". The AC2 agreement coefficient for specific steps was, in some instances, higher in the first step of the tool than in the second. Results on step-level ranged from 0.43 (95% CI: 0.24, 0.62, classified as "fair or higher") in the core item 3 step 2 to 0.78 (95% CI: 0.79, 0.92, classified as "almost perfect") in the core item 1 step 1.
CONCLUSION
The results support the perspective that the ROBUST-RCT is a reliable and straightforward tool for assessing the risk of bias in systematic reviews. Taken together with previous findings, the higher agreement on some items in the first step may support the view that authors of future systematic reviews should transparently report both steps, enabling readers to build their own reasoning from ratings in step 1.
PLAIN LANGUAGE SUMMARY
Risk of bias tools are instruments used in the synthesis of medical scientific literature to assess whether specific characteristics of clinical trials could affect their results. The ROBUST-RCT is one of these tools and was recently introduced with characteristics that may enable junior researchers to conduct those assessments. This study evaluates the tool's inter-rater reliability, the extent to which users agree in their assessments. A result of 0.00 would mean no better agreement than random data, while 1.00 would suggest perfect agreement - an ideal not often met. In this study, the junior researchers had an inter-rater reliability of 0.59 for the most relevant step of the tool, which is interpreted as at least moderate agreement. Therefore, it supports that the risk of bias could be assessed by junior scientists using the ROBUST-RCT tool.
P. Vidor, Sofia Simoni Rossi Fermo, Y. Casiraghi et al.· Journal of Clinical Epidemio...· 0 citations
Abstract Objectives To compare the reliability of four commonly used structured tools for risk-of-bias assessment in prognostic cohort studies and evaluate their agreement with expert appraisal. Design Methodological comparison study. Setting Secondary analysis of published prognostic cohort studies in patients with acute pulmonary embolism from a previously conducted systematic review. Participants Sixty-three cohort studies assessing the prognostic role of echocardiography in patients with acute pulmonary embolism. Primary and secondary outcome measures Four independent reviewers assessed study quality/risk of bias using the Newcastle-Ottawa Scale (NOS), Risk Of Bias in Non-Randomised Studies of Interventions (ROBINS-I), Quality In Prognostic Studies (QUIPS) and Quality Assessment of Diagnostic Accuracy Studies (QUADAS-2). Concomitantly and independently, expert reviewers’ implicit global judgements were used as an external comparator. Inter-rater reliability, agreement with expert appraisal, internal consistency and floor/ceiling effects were assessed. Results Inter-rater agreement was fair for NOS (Gwet’s agreement coefficient 1, 0.25, 95% CI 0.15 to 0.34), ROBINS-I (0.38, 95% CI 0.28 to 0.47) and QUADAS-2 (0.35, 95% CI 0.27 to 0.43) and almost perfect for QUIPS (0.85, 95% CI 0.73 to 0.93). Agreement with expert appraisal was limited for all tools except ROBINS-I, which showed fair concordance. QUIPS frequently classified studies as low risk of bias, suggesting potential overestimation of study quality. Internal consistency was generally low across tools, while ceiling effects were observed for NOS and QUIPS. ROBINS-I showed the most balanced distribution of ratings. Conclusions Structured tools for risk-of-bias assessment in prognostic cohort studies have variable reliability and limited agreement with expert appraisal. ROBINS-I showed the strongest concordance with expert judgement, whereas QUIPS may provide optimistic ratings. These findings support further refinement and standardisation of risk-of-bias assessment methods for observational prognostic research.
L. Cimini, F. Klok, M. Carrier et al.· BMJ Open· 0 citations
It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.
Euijun Yang, S. Ko, Hyekyung Woo· Journal of Medical Internet...· 0 citations
BACKGROUND
This article aims to provide an overview of the risk of bias, concerns regarding applicability, and the reporting of findings from studies examining the diagnostic test accuracy (DTA) of self-report questionnaires used to screen for anxiety disorders.
METHODS
We merged data from 103 studies included in six systematic reviews on the diagnostic accuracy of eight anxiety questionnaires (HADS-A, GAD-7, GAD-2, BAI, STAI-S, STAI-T, PROMIS-A-SF8a, and OASIS) in adults. To be included, primary studies had to report the sensitivity and specificity for at least one cut-off when tested against a (semi-)structured clinical interview used as a reference standard. Risk of bias and applicability were rated with the QUADAS-2 tool (Quality Assessment of Diagnostic Accuracy Studies).
RESULTS
The 103 studies involved a total of 34,833 participants (range 48 to 3215). The most commonly used questionnaires were the HADS-A (54 studies), the GAD-7 (46), and the GAD-2 (31). Overall, 80% of the studies were rated as having unclear or high risk of bias, and for 88%, concerns regarding applicability were rated as unclear or high. Uncertainties or concerns were particularly common regarding participant recruitment and whether the samples included reflected realistic screening populations. Publications often reported sensitivity and specificity in a very selective manner for a single or a few cut-offs only.
CONCLUSIONS
Most DTA studies on questionnaires used to screen for anxiety have significant shortcomings in the reporting of methods and results. Our findings provide a basis for recommendations for future studies.
Klaus Linde, Alexey Fomenko, Zekeriya Aktürk et al.· Journal of Psychosomatic Res...· 0 citations
Abstract Background A growing body of literature leverages large language models (LLMs) to make mental health predictions. However, these models are prone to bias, and studies to validate their clinical utility are lacking. Objective This scoping review aims to uncover bias and clinical utility limitations stemming from the methodological design of LLM-based mental health predictive systems. In addition, it intends to document the level of self-reflection about bias and clinical challenges reported by authors in their own work. Methods This work follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and was registered online. Eligible studies were original research articles in English published between 2019 and 2024, using LLMs to detect mental health conditions in nonsynthetic textual data. The search was conducted in 5 scientific databases (PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and ACL Anthology) with queries associating keywords related to “Mental Health,” “Large Language Models,” and “Prediction.” We extracted both methodological information about the included studies and authors’ statements relevant to issues of bias and clinical utility. This extraction was based on a framework screening the entire pipeline of development of LLMs with applications in mental health: research design and selection, data collection, outcome definition, model development, and postdeployment considerations. Statistical description of the retrieved entities, as well as thematic coding, was performed for analysis. Results A total of 2472 articles were identified, of which 263 (10.6%) were assessed for eligibility, and 201 (8.1%) were included in the review. Included studies were mostly recent, indicating a growing interest in the use of LLMs for mental health predictions. Our analysis revealed that a majority of studies share similar methodological choices along their development pipeline: most of them focus on depressive disorders identified via processing user texts on social media, mainly with the use of nonspecialist LLMs derived from BERT (Bidirectional Encoder Representations from Transformers). Following previous works on these matters, we highlighted how these choices may hinder the clinical relevance and fairness of the envisioned systems. Similarly, we found that 164 (81.6%) studies mention themes related to bias and clinical utility; however, most of the discussion revolves around data-centered issues. Only 41 (20.4%) articles mention themes associated with at least 3 out of 5 pipeline steps, suggesting a limited appropriation of the notions of bias and clinical utility in such a sensitive context as mental health analysis. Conclusions Bias and clinical utility are lightly covered in the field of LLM-based mental health prediction research as of 2019‐2024. In-depth approaches involving interdisciplinary teams of clinicians and natural language processing specialists are needed to ensure technical soundness, clinical relevance, and fair outcomes for potential users.
Clémentine Bleuze, Karen Fort, Vincent P. Martin et al.· JMIR AI· 0 citations