Evaluating LLMs for Automated Extraction and Classification of Science, Technology and Innovation (STI) Indicators
Abstract
Science, Technology and Innovation (STI) indicators are used to measure scientific capacity, technological development, and innovation performance. Collecting these indicators from the scientific literature is slow and difficult to scale. This study evaluates two large language models (LLMs)—Gemini and Claude—on automatic extraction and classification of STI indicators from five peer-reviewed articles. A Seed Prompt strategy—in which a model generates its own operational prompt from a set of predefined criteria—inspired by automatic prompt engineering (Zhou et al., 2022), was used. Model outputs were compared against a Gold Standard—a reference list of indicators manually identified by human experts—using a strict Exact Match criterion, in which an extracted indicator is accepted only if its normalized string is identical to one in the Gold Standard. Extraction F1 averaged 0.265 for Gemini and 0.356 for Claude, with high variation across documents (0.0–0.977). Mean Adjusted Rand Index (ARI) was 0.050 for Claude and 0.119 for Gemini, indicating near-random classification. Approximately 69% of indicators were left unclassified. Performance was higher for lexically standardized indicator categories and lower for categories with unstable definitions. These results show that current LLMs are not reliable for unsupervised STI indicator workflows. Hybrid human-AI approaches are recommended as a near-term alternative.