Skip to content

BELXTR: Biomedical Entity Linking via Contextualized Token Retrieval

Sep 2026 · 0 citations · 81 references
Computer Science

TL;DR

BELXTR is presented, a novel embedding model based on the multi-vector (a.k.a. late interaction) architecture, which allows to leverage token-level matching information in biomedical entity linking by integrating an existing task-specific training objective and exploring active query expansion.

Abstract

Biomedical Entity Linking disambiguates mentions to entities in a knowledge base (KB), making it the cornerstone of information extraction pipelines. While embedding-based models are a popular approach for the task, they suffer from a key limitation. They compress mentions (and entities) into a single vector, forcing the model to average away crucial fine-grained differences. We present BELXTR, a novel embedding model based on the multi-vector (a.k.a. late interaction) architecture, which allows to leverage token-level matching information. BELXTR extends the original XTR model to biomedical entity linking by integrating an existing task-specific training objective and exploring active query expansion. Experiments across ten corpora and five KBs show that BELXTR improves upon current state-of-the-art in half of the corpora with an average improvement of 5pp recall@1. The largest gains are reported on the challenging cross-species gene disambiguation subtask, where BELXTR outperforms an LLM-powered retrieve-and-rerank pipeline and closely approaches a specialized rule-based system. Our results highlight multi-vector models as a practical alternative to hard-to-maintain rule-based systems or in scenarios where LLM-based reranking is too costly as in PubMed-scale mining. The code to reproduce our experiments can be found at: https://github.com/sg-wbi/belxtr.

View source

Similar papers

#small language model Open access Sep 2026

Biomedical retrieval-augmented generation for relation classification

The rapid expansion of biomedical literature requires automated methods for accurate and efficient information extraction. This study addresses relation classification: given a pair of annotated biomedical entities in a research article title and abstract, assigning the relation that holds between them from a pre-defin...

Jannat, Charlie Dil, Tom Arodz et al. · 0 citations
Open access Aug 2026

Biomedical Text Mining and Information Extraction Using Prompt-Enhanced and LoRA-Adapted Large Language Models

Biomedical named entity recognition (NER) and relation extraction (RE) remain challenging because biomedical texts contain ambiguous abbreviations, complex entity boundaries, domain-specific terminology, and implicit relations. This study proposes a prompt-enhanced and QLoRA-adapted large language model framework for b...

Feng Yan, De-Quan Zheng, Feng Yu et al. · 0 citations

Addressing Span Imbalance and Semantic Complexity in Nested Medical Named Entity Recognition

As a fundamental task in biomedical natural language processing, Medical Named Entity Recognition (MNER) aims to identify and classify medical entities from unstructured medical texts. A major challenge in this task is the prevalence of nested entities, which arise from the syntactic complexity and domain-specific char...

Yuling Li, Yang-Juan Hu, Yi-Ming Bao et al. · 0 citations
#natural language process... Preprint Sep 2026

Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking

A simple, training-free framework in which a reasoning-capable vision-language model iteratively searches and reasons over Wikipedia, gathering evidence dynamically, shows that reasoning and retrieval are complementary on rare entities.

Parinthapat Pengpun, Simran Khanuja, Graham Neubig · 0 citations
Open access Aug 2026

Language-Model-Based Architecture for Automatic Concept Placement in Ontologies

This paper addresses the placement of concepts that are absent from the target ontology—the out-of-knowledge-base setting—in which a textual mention must be assigned one or more insertion positions in the subsumption hierarchy rather than linked to an existing node.

Zhanna B. Sadirmekova, M. Sambetbayeva, B. Abdygalym et al. · 0 citations
#natural language process... Preprint Aug 2026

Bridging Lexical Divergence: LLM-Assisted, Cost-Efficient, Zero-shot Scientific Entity Linking

The Sci-ZSEL framework is proposed, a framework that selectively generates entity aliases with an LLM to control computational cost, and applies an ontology-aware filter to remove aliases that semantically drift toward ontology neighbors.

Md Rasel Khondokar, Qiao Qiao, Farjana Sultana Samia et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.