ENTGPT: Applying Large Language Models to Systematic Review Screening With the Novel STARR Protocol.
OBJECTIVES Systematic literature reviews (SLRs) are time-intensive and resource-consuming. While large language models (LLMs) have shown promise in encoding clinical knowledge, evidence for their performance on complex text analysis, necessary for SLRs in otolaryngology, remains limited. In this proof-of-concept study, we investigate an LLM's performance in screening relevant articles for a SLR with the novel screening of title and abstracts, reevaluation, and full-text review (STARR) protocol. METHODS ENTGPT (based on GPT-4o) was compared to two human reviewers in article inclusion/exclusion decisions using the traditional and STARR screening protocols. ENTGPT was provided with inclusion and exclusion criteria, titles, abstracts, and full texts (if available) of the 850 articles retrieved in the original search. The model's decisions were compared to those made by two human reviewers. RESULTS ENTGPT, using the STARR protocol, achieved 99.87% accuracy in article classification compared with human reviewers (95% CI: 0.99-1.0), including 100% specificity and 95% sensitivity. When using the traditional protocol, sensitivity declined to 35%. ENTGPT, using the traditional protocol, achieved 99.47% accuracy in article classification compared with human reviewers (95% CI: 0.99-1.0), including 100% specificity and 35% sensitivity. When using the STARR protocol, accuracy improved to 99.87% and sensitivity markedly increased to 95%. CONCLUSIONS ENTGPT accurately replicated human reviewers in article selection and data extraction for an otolaryngology SLR using the STARR and traditional protocols. This performance suggests that LLMs could be employed to significantly streamline the SLR process, potentially saving substantial time and resources for researchers. LEVEL OF EVIDENCE N/A.