Skip to content

Knowledge-enhanced LLMs for multilingual biomedical concept normalization: a multilingual benchmarking and behavioral analysis

Sep 2026 · npj Digital Medicine · 0 citations
Biomedical Text Mining and Ontologies

TL;DR

It is demonstrated that structured knowledge enrichment is critical for effective LLM-based multilingual concept normalization, while surface-form sensitivity and positional biases remain important challenges for fully automated clinical pipelines.

Abstract

Biomedical concept normalization (BCN)—mapping clinical terms to standardized concepts in biomedical terminologies—is a key bottleneck for semantic interoperability across health information systems, particularly in multilingual settings and when terms appear with minimal surrounding context, as is common in coded fields of real-world electronic health record databases. Large language models (LLMs) offer new opportunities to automate BCN, yet their behavior under these conditions remains poorly understood. This study introduces MedLexAlign, a multilingual benchmark comprising 52,011 unique terms mapped to 22,787 biomedical concepts across five European languages (English, French, German, Spanish, and Turkish). It proposes a knowledge-enhanced retrieve-then-rerank pipeline employing discriminative LLMs as dense retrievers and generative LLMs as rerankers, augmented with structured knowledge from the Unified Medical Language System (UMLS). We evaluate eight discriminative LLMs as dense retrievers and six generative LLMs (8B–32B parameters; across different reasoning modes) as rerankers, examining abbreviation handling, synonym sensitivity, and positional bias. The best-performing dense retriever, nemotron, achieved R@1 of 0.57 and R@10 of 0.80, outperforming kalm-gemma (R@1 = 0.56, R@10 = 0.78; p -value = 0.02). For reranking, UMLS-based knowledge enrichment yielded consistent gains: qwen3-32b-r achieved ΔR@1 of +0.10 with enrichment versus +0.05 without ( p -value < 0.001). However, LLMs exhibited sensitivity to lexical surface forms, handled synonymous representations inconsistently, and showed strong primacy bias toward earlier-ranked candidates. These findings demonstrate that structured knowledge enrichment is critical for effective LLM-based multilingual concept normalization, while surface-form sensitivity and positional biases remain important challenges for fully automated clinical pipelines.

Read PDF

Similar papers

Open access Sep 2026

A Comparative Benchmark of Biomedical Language Models for Concept Normalization from Real-World Text

A benchmark-guided, scalable framework for automated medical terminology standardization that accepts heterogeneous short medical expressions without manual input pre-processing and automatically performs text refinement, semantic retrieval and terminology mapping to standardized concepts and vocabulary codes is establ...

Anshul Verma, Abhijay, Manan Vangani et al. · 0 citations
Open access Aug 2026

A unified framework and benchmark for generalizable biomedical knowledge extraction and applications with large language models

Results demonstrate that InfoFlowEX equips LLMs with robust adaptability, achieving consistent gains over baselines with minimal task-specific customization, highlighting InfoFlowEX for real-world biomedical applications.

Wuyang Lan, Siqi Zhang, Wenzheng Wang et al. · 0 citations
#large language models Open access Sep 2026

Medical Concept Normalization of German Clinical Expressions to SNOMED CT.

INTRODUCTION Clinical narratives in electronic health records frequently contain clinical expressions describing medical conditions. Their free-text format limits interoperability and automated processing. Medical concept normalization (MCN) addresses this challenge by mapping textual expressions to standardized termin...

Helena Adam, Akhila Abdulnazar, Roland Roller et al. · 0 citations
#natural language process... Preprint Aug 2026

En-ViMedNER: An English-Vietnamese Parallel Biomedical Corpus with UMLS Semantic Type Annotations

En-ViMedNER is presented, the first English-Vietnamese parallel biomedical NER corpus annotated with UMLS semantic types, which are language-neutral codes providing a shared cross-lingual label space and ensuring direct comparability with existing UMLS-based resources.

Nhu Vo, P. Nguyen, Nu-Uyen-Phuong Le et al. · 0 citations
#large language models Review Open access Sep 2026

Large Language Models for Clinical Note Simplification: A Systematic Review and Experimental Evaluation of Medical Text Readability.

The findings suggest that conventional readability metrics should be extended with domain-specific measures to more accurately assess comprehensibility in medical texts and that large Language Models show strong potential to enhance the accessibility of clinical documentation for patients.

M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al. · 0 citations
Preprint Aug 2026

Analyzing and Mitigating Cross-Lingual Degradation in Multilingual Medical VQA

A multilingual medical VQA benchmark over eight languages is constructed, organized into four representative scenarios that isolate the core capabilities medical VQA requires, and a training-free scenario-aware representation engineering method is proposed, leveraging LVLMs's superior English medical VQA capability to...

Jingbo Wang, Sendong Zhao, Haochun Wang et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.