Free-text clinical records represent an untapped wealth of data for secondary use, but realising their potential is limited by resource demands necessary for accurate information extraction at scale. We introduce a scalable, resource-efficient, and high-performance information extraction pipeline that leverages large language models (LLMs) to address these challenges. Our pipeline was developed and tested using real-world dual specialist-annotated ophthalmic clinical letters, and achieved strong performance with a proprietary model in development, yielding a maximum micro-averaged F1 score of 0.954 (95% CI 0.941–0.967) for diagnosis across nine conditions through iterative prompt refinement alone, also demonstrating strong generalisability (micro-F1 0.945–0.980) in temporal validation. This approach was extended to other models in the same family and 17 LLMs from seven open-weight LLM families. Beyond performance, we develop a multi-dimensional assessment for deployment in data extraction tasks, including an error taxonomy and Pareto frontier analyses to systematically map the operational trade-offs across different LLM configurations. A robust approach to operationalisation in real-world workflows at scale may help lay the foundation for next-generation data pipelines that accelerate scientific discovery and power continuous learning health systems.
A. Y. Ong, Quang Nguyen, I. Barai et al.· npj Digital Medicine· 1 citation
While AI agents delivered highly efficient, directionally aligned assessments, they did not fully capture the nuances of human clinical judgment and could not substitute for physician-centered evaluation and promise assistive tools that can triage or pre-screen outputs to reduce human burden.
Peilun Shi, Jian Li, Ziqi Yang et al.· npj Digital Medicine· 0 citations