Skip to content
Review Open access

The ROBUST-RCT tool had at least moderate agreement in an inter-rater reliability study of assessments by junior researchers.

Jul 2026 · Journal of Clinical Epidemiology · pp. 112418 · 0 citations
Medicine

Abstract

Objective

The recently introduced ROBUST-RCT tool aims to reconcile ease of application with methodological rigor in risk-of-bias assessments for systematic reviews. The tool is structured in two steps: first, evaluating what happened; second, judging the risk of bias related to the assessed aspect of the study. Its straightforward design aims to avoid overly complex workflows. Thus, its usability testing included junior reviewers to ensure accessibility and ease of use. No data regarding its inter-rater reliability are currently available. This study aims to assess the inter-rater reliability of the ratings between junior researchers using the ROBUST-RCT. STUDY

Design

An inter-rater reliability study. Four junior researchers screened and rated a random sample of 115 articles from a systematic search on PubMed. An additional phase beyond the prospectively defined research phases was introduced to exclude articles in which the two raters assessed different outcomes, resulting in a sample of 85 articles. As pre-specified in the protocol, the primary statistical analysis employed Gwet's AC2 at each step for each core item ("step-level") and at aggregated step 2 ratings ("judgment set") to provide an overview of the final judgment in the tool. Exploratory analyses include Fleiss' Kappa and a block-level approach.

Setting

Universidade Federal do Rio Grande do Sul, a university in southern Brazil.

Results

In the primary analysis, the aggregated data with the step 2 ratings ("judgment set") yielded a Gwet's AC2 agreement coefficient of 0.59 (95% CI: 0.53, 0.65); its inter-rater reliability was classified in Gwet's benchmarking as "moderate or higher". The AC2 agreement coefficient for specific steps was, in some instances, higher in the first step of the tool than in the second. Results on step-level ranged from 0.43 (95% CI: 0.24, 0.62, classified as "fair or higher") in the core item 3 step 2 to 0.78 (95% CI: 0.79, 0.92, classified as "almost perfect") in the core item 1 step 1.

Conclusion

The results support the perspective that the ROBUST-RCT is a reliable and straightforward tool for assessing the risk of bias in systematic reviews. Taken together with previous findings, the higher agreement on some items in the first step may support the view that authors of future systematic reviews should transparently report both steps, enabling readers to build their own reasoning from ratings in step 1. PLAIN LANGUAGE SUMMARY Risk of bias tools are instruments used in the synthesis of medical scientific literature to assess whether specific characteristics of clinical trials could affect their results. The ROBUST-RCT is one of these tools and was recently introduced with characteristics that may enable junior researchers to conduct those assessments. This study evaluates the tool's inter-rater reliability, the extent to which users agree in their assessments. A result of 0.00 would mean no better agreement than random data, while 1.00 would suggest perfect agreement - an ideal not often met. In this study, the junior researchers had an inter-rater reliability of 0.59 for the most relevant step of the tool, which is interpreted as at least moderate agreement. Therefore, it supports that the risk of bias could be assessed by junior scientists using the ROBUST-RCT tool.

Read PDF