A benchmark-guided, scalable framework for automated medical terminology standardization that accepts heterogeneous short medical expressions without manual input pre-processing and automatically performs text refinement, semantic retrieval and terminology mapping to standardized concepts and vocabulary codes is established.
Abstract
Patient-reported and clinically documented narratives often contain informal, fragmented and linguistically heterogeneous expressions, complicating medical concept normalization (MCN) and frequently necessitating pre-processing before terminology mapping. Despite rapid advances in biomedical language modeling, the comparative utility of representation models for MCN and their integration with instruction-tuned LLMs in automated normalization workflows remain underexplored. In this study, we systematically benchmarked 15 general, biomedical and clinical transformer-based representation models together with 12 instruction-tuned LLMs across 12,713 instances from five established datasets: TAC2017_ADR, TwADR-L, TwiMed, CADEC and SMM4H2017. Representation models were evaluated using embedding-based semantic retrieval, whereas instruction-tuned LLMs were assessed as upstream text-correction modules. SapBERT achieved the highest Top-5 accuracy among representation models, reaching 63.8% for SNOMED CT and 58.0% for MedDRA without correction. Qwen 2 Instruct was selected as the preferred corrector on the basis of its favorable balance between Top-1 normalization performance and computational efficiency relative to substantially larger models, including Llama 3.1 Instruct (70B). Incorporation of Qwen 2 Instruct increased Top-5 accuracy to 69.4% for SNOMED CT and 63.2% for MedDRA. The resulting framework accepts heterogeneous short medical expressions without manual input pre-processing and automatically performs text refinement, semantic retrieval and terminology mapping to standardized concepts and vocabulary codes. These findings establish a benchmark-guided, scalable framework for automated medical terminology standardization.
INTRODUCTION
Clinical narratives in electronic health records frequently contain clinical expressions describing medical conditions. Their free-text format limits interoperability and automated processing. Medical concept normalization (MCN) addresses this challenge by mapping textual expressions to standardized termin...
Helena Adam, Akhila Abdulnazar, Roland Roller et al.· Studies in Health Technology...· 0 citations
It is demonstrated that structured knowledge enrichment is critical for effective LLM-based multilingual concept normalization, while surface-form sensitivity and positional biases remain important challenges for fully automated clinical pipelines.
H. Rouhizadeh, A. Yazdani, Boya Zhang et al.· npj Digital Medicine· 0 citations
Results demonstrate that InfoFlowEX equips LLMs with robust adaptability, achieving consistent gains over baselines with minimal task-specific customization, highlighting InfoFlowEX for real-world biomedical applications.
Wuyang Lan, Siqi Zhang, Wenzheng Wang et al.· Cell Reports Medicine· 0 citations
Findings demonstrate that compact models can achieve strong biomedical classification performance through KD under compatible teacher–student pairings, while also highlighting that KD effectiveness varies substantially depending on the specific model combination.
The findings suggest that conventional readability metrics should be extended with domain-specific measures to more accurately assess comprehensibility in medical texts and that large Language Models show strong potential to enhance the accessibility of clinical documentation for patients.
M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al.· Studies in Health Technology...· 0 citations
It is demonstrated that modern LLMs with reasoning capabilities can effectively support real-time clinical evidence synthesis without task-specific fine-tuning, offering a pathway toward scalable automated systems for clinical trial interpretation that could substantially reduce the evidence-to-practice gap in medical...
Shobanapriyan Chandrasegaran, Amal Htait· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.