USING LARGE LANGUAGE MODELS FOR LITERATURE SEARCH IN CARDIOVASCULAR SURGERY SYSTEMATIC REVIEWS AND META-ANALYSES
Abstract
Highlights Large Language Models (LLMs) enable automation of initial abstract screening in systematic reviews, significantly reducing manual workload. The effectiveness of LLMs heavily depends on prompt engineering, which must clearly translate inclusion and exclusion criteria into actionable instructions. Two prompt design strategies are presented and illustrated: a “lenient” approach (maximizing recall) and a “strict” approach (reducing false positives). A practical workflow for downloading and processing PubMed abstracts using LLMs (e.g., DeepSeek, GPT‑4o, Perplexity) is demonstrated. Key limitations of LLMs – hallucinations, sensitivity to phrasing, and training‑data biases — are discussed, emphasizing the need for expert validation. A two‑stage screening strategy is recommended: first a lenient pass, followed by a stricter filter, with subsequent manual verification of relevant articles. Prompt engineering is framed as an iterative, domain‑specific art rather than a one‑size‑fits‑all procedure, requiring continuous refinement for each research question. Abstract Large Language Models have evolved into powerful tools for automating the primary screening of articles’ abstracts in systematic reviews, enabling a significant reduction in manual labor. The article presents a comprehensive review of prompt engineering principles, demonstrating how traditional meta-analysis criteria can be transformed into clear instructions for artificial intelligence. Using case studies in coronary artery bypass grafting and congenital heart disease surgery, we illustrate the impact of prompt formulation on the comprehensiveness, accuracy, and overall efficiency of literature screening. Furthermore, typical errors are discussed, and the ongoing necessity of expert oversight to minimize hallucinations and biases inherent in artificial intelligence conclusions is emphasized. Ultimately, systematic prompt engineering combined with expert evaluation allows researchers in cardiovascular surgery to optimize the search for meta-analysis sources based on Large Language Models, ensuring a faster and more comprehensive synthesis of evidence and methods for processing primary material.