Skip to content
Open access

A Comparative Benchmark of Biomedical Language Models for Concept Normalization from Real-World Text

Sep 2026 · bioRxiv · 0 citations · 7 references
Biology

TL;DR

A benchmark-guided, scalable framework for automated medical terminology standardization that accepts heterogeneous short medical expressions without manual input pre-processing and automatically performs text refinement, semantic retrieval and terminology mapping to standardized concepts and vocabulary codes is established.

Abstract

Patient-reported and clinically documented narratives often contain informal, fragmented and linguistically heterogeneous expressions, complicating medical concept normalization (MCN) and frequently necessitating pre-processing before terminology mapping. Despite rapid advances in biomedical language modeling, the comparative utility of representation models for MCN and their integration with instruction-tuned LLMs in automated normalization workflows remain underexplored. In this study, we systematically benchmarked 15 general, biomedical and clinical transformer-based representation models together with 12 instruction-tuned LLMs across 12,713 instances from five established datasets: TAC2017_ADR, TwADR-L, TwiMed, CADEC and SMM4H2017. Representation models were evaluated using embedding-based semantic retrieval, whereas instruction-tuned LLMs were assessed as upstream text-correction modules. SapBERT achieved the highest Top-5 accuracy among representation models, reaching 63.8% for SNOMED CT and 58.0% for MedDRA without correction. Qwen 2 Instruct was selected as the preferred corrector on the basis of its favorable balance between Top-1 normalization performance and computational efficiency relative to substantially larger models, including Llama 3.1 Instruct (70B). Incorporation of Qwen 2 Instruct increased Top-5 accuracy to 69.4% for SNOMED CT and 63.2% for MedDRA. The resulting framework accepts heterogeneous short medical expressions without manual input pre-processing and automatically performs text refinement, semantic retrieval and terminology mapping to standardized concepts and vocabulary codes. These findings establish a benchmark-guided, scalable framework for automated medical terminology standardization.

Read PDF

Similar papers

#large language models Open access Sep 2026

Medical Concept Normalization of German Clinical Expressions to SNOMED CT.

INTRODUCTION Clinical narratives in electronic health records frequently contain clinical expressions describing medical conditions. Their free-text format limits interoperability and automated processing. Medical concept normalization (MCN) addresses this challenge by mapping textual expressions to standardized termin...

Helena Adam, Akhila Abdulnazar, Roland Roller et al. · 0 citations
#large language models Open access Sep 2026

Knowledge-enhanced LLMs for multilingual biomedical concept normalization: a multilingual benchmarking and behavioral analysis

It is demonstrated that structured knowledge enrichment is critical for effective LLM-based multilingual concept normalization, while surface-form sensitivity and positional biases remain important challenges for fully automated clinical pipelines.

H. Rouhizadeh, A. Yazdani, Boya Zhang et al. · 0 citations
Open access Aug 2026

A unified framework and benchmark for generalizable biomedical knowledge extraction and applications with large language models

Results demonstrate that InfoFlowEX equips LLMs with robust adaptability, achieving consistent gains over baselines with minimal task-specific customization, highlighting InfoFlowEX for real-world biomedical applications.

Wuyang Lan, Siqi Zhang, Wenzheng Wang et al. · 0 citations
Open access 2026

Knowledge Distillation for Biomedical Text Classification: A Systematic Comparative Analysis of Multiple Teacher–Student Architectures

Findings demonstrate that compact models can achieve strong biomedical classification performance through KD under compatible teacher–student pairings, while also highlighting that KD effectiveness varies substantially depending on the specific model combination.

Amine Gonca Toprak, Aytuğ Onan · 0 citations
#large language models Review Open access Sep 2026

Large Language Models for Clinical Note Simplification: A Systematic Review and Experimental Evaluation of Medical Text Readability.

The findings suggest that conventional readability metrics should be extended with domain-specific measures to more accurately assess comprehensibility in medical texts and that large Language Models show strong potential to enhance the accessibility of clinical documentation for patients.

M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al. · 0 citations
Review 2026

Enhancing Clinical Trial Analysis through Large Language Models for Multi-Evidence Natural Language Inference

It is demonstrated that modern LLMs with reasoning capabilities can effectively support real-time clinical evidence synthesis without task-specific fine-tuning, offering a pathway toward scalable automated systems for clinical trial interpretation that could substantially reduce the evidence-to-practice gap in medical...

Shobanapriyan Chandrasegaran, Amal Htait · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.