Accelerating metadata annotation in collaborative research centers: A hybrid AI workflow for biomedical entities
Abstract
Collaborative Research Centers rely on FAIR-compliant, richly structured metadata, yet manual annotation is a major bottleneck. We implemented a search-augmented large language model (LLM) workflow within a local research data management system to pre-annotate biomedical entities, using human-in-the-loop verification to ensure data quality. The pipeline uses Gemini 3 Pro in a two-step prompting strategy: (1) identify dataset deposits and stable identifiers in articles converted to Markdown; (2) extract structured fields from curated repository landing pages rendered via a headless browser. To handle a highly hierarchical metadata schema, we flattened the schema for prompting and remapped outputs to strict JSON with granular provenance tags. Authors received pre-filled metadata and could accept, edit, or delete entries (TP, FP, FN mapping). Performance metrics (precision, recall, F1) were estimated as proportions and synthesized via random-effects meta-analysis. Among 51 screened articles, the LLM identified a repository deposit in 31; authors responded for 17 (55% of articles with an identified deposit; 33% of all screened articles), yielding 39 verified datasets. False positives were rare (mean 0.23) and false negatives low (mean 1.46). Precision was consistently high at 99.65% (95% CI 98.42%–100.00%). Recall showed more variability at 93.75% (95% CI 89.79%–96.96%), and the F1 score reached 96.17% (95% CI 93.78%–98.11%). The hybrid workflow achieved high author-verified precision with moderate recall, shifting effort from drafting to reviewing while maintaining schema compliance. As these metrics rely on author acceptance and response rates were modest, broader, multi-site validation is needed.