Skip to content
Open access

TrialCode Agent: LLM-Assisted Clinical Code-Set Construction for Trial Emulation

Aug 2026 · medRxiv · 0 citations
Medicine

TL;DR

Combining LLM-based clinical interpretation with deterministic terminology verification produces auditable, database-ready code sets, but retrieval and broad family expansion do not consistently improve exact-code performance.

Abstract

Objective Translating free-text clinical trial criteria into computable code sets is a valuable standardization practice that is necessary for producing reproducible real-world evidence studies but requires standardized interpretation across multiple clinical vocabularies. Methods We developed TrialCode Agent, a hybrid-large language model (LLM)-terminology verification agent that generates, formats, verifies, and expands candidate codes from free-text clinical criteria. The system supports ICD-9-CM diagnoses and procedures, ICD-10-CM, ICD-10-PCS, LOINC, and RxNorm medication concepts. We compared Baseline, Hybrid biomedical retrieval-augmented generation (RAG), and terminology-guided Family expansion pipelines using Claude, GPT Qwen, and MedGemma on 40 criteria from 11 trial groups. Performance was evaluated against expert-built reference code sets using exact-code precision, recall, and F1. Results The optimal pipeline varied by model. Claude with Baseline achieved the highest performance (precision 0.755, recall 0.619, F1 0.680), followed by GPT-5.5 with Baseline (precision 0.569, recall 0.658, F1 0.610), Qwen with Hybrid biomedical RAG (precision 0.656, recall 0.470, F1 0.548), and MedGemma with Family expansion (precision 0.487, recall 0.316, F1 0.383). Hybrid biomedical RAG improved aggregate F1 only for Qwen but increased GPT-5.5 RxNorm F1 from 0.320 to 0.909. Macro-averaged results showed criterion-level gains despite lower micro-averaged aggregate performance. Family expansion increased recall across models but generally reduced precision. In staged verifier ablation, micro-F1 increased from 0.254 before verification to 0.505 after final verification and expansion. Existence/vocabulary checking removed 2,594 false-positive codes, and acceptance filtering removed 952 additional false-positive codes before controlled expansion. Conclusions Combining LLM-based clinical interpretation with deterministic terminology verification produces auditable, database-ready code sets, but retrieval and broad family expansion do not consistently improve exact-code performance. Retrieval was particularly useful for RxNorm mapping, whereas overly broad or incomplete candidate generation remained the main source of error. Deterministic verification improves code validity and query readiness but cannot replace accurate clinical interpretation.

Read PDF

Similar papers

Open access Sep 2026

Clinical Code Mapping with LLM Tool Use: A Pilot for Automated Data Extraction of Medication and Diagnosis Information from Unstructured Clinical Notes.

LLMs are suitable for information extraction of medications from clinical notes for use in research databases, however, for a clinical setting where the treatment of patients would be dependent on LLM performance, the current state-of-the-art open weight models are not accurate enough.

T. Spreuer, A. Günther, R. Majeed · 0 citations
Review Open access Sep 2026

Automated Extraction of Genetic Eligibility Criteria from Clinical Trial Records Using LLMs - A Technical Case Report.

INTRODUCTION Accurate interpretation of clinical trial eligibility criteria is essential for applications such as patient-trial matching and clinical decision support, particularly in precision oncology. However, relevant information, including genetic mutation requirements, is typically embedded in unstructured text w...

Georg Mathes, S. Berger, Stefan Sigle · 0 citations
#artificial intelligence Preprint Sep 2026

Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation

Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL), retain the execut...

Erfan D. Dehkalani, S. Shankaran, Abbot R. Laptook et al. · 0 citations
#natural language process... Preprint Aug 2026

KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs

This work proposes Knowledge-Guided Reasoning over Clinical Evidence with LLMs (KREL), a framework that leverages LLMs for clinical text understanding and reasoning while integrating external ICD coding guidelines as structured knowledge, and enables tight coupling between domain knowledge and LLM reasoning.

Xubin Chen, Yi-Peng Zhou, Wenxin Sun et al. · 0 citations
Open access Aug 2026

Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals

Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.

L. Barrett, N. Joshi, A. S. North et al. · 0 citations
Preprint Aug 2026

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower, and this results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves gro...

P. Reddy, C. Mandke, Suvrankar Datta et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.