Skip to content
Open access

Systematic assessment of text summarization methods for biomedical literature from frequency methods to language models

Sep 2026 · iScience · Vol 29 · 0 citations · 72 references
Medicine

TL;DR

These findings provide a systematic reference for selecting biomedical summarization tools and highlight that broad pretraining outperforms narrow domain adaptation.

Abstract

Summary The rapid expansion of biomedical literature demands automated summarization tools that reliably condense research articles into concise, accurate summaries. We benchmarked 62 summarization methods, ranging from frequency-based and TextRank extractors to encoder-decoder models (EDMs) and large language models (LLMs), on 1,000 biomedical abstracts from 20 journals across ScienceDirect and Cell Press, using author-written highlights as reference summaries. Models were evaluated with a composite suite of lexical, semantic, and factual metrics, including ROUGE, BLEU, METEOR, embedding-based similarity, and factuality scores. General-purpose models (e.g., Mistral, GPT, and Llama) achieved the highest overall performance across lexical and semantic dimensions, outperforming reasoning-oriented (e.g., DeepSeek and Magistral) and domain-specific (e.g., BioGPT and BioMistral) models. Notably, medium-sized models outperformed large-scale models, suggesting an optimal balance between model capacity and efficiency, while classical extractive methods lagged behind neural approaches. These findings provide a systematic reference for selecting biomedical summarization tools and highlight that broad pretraining outperforms narrow domain adaptation.

Read PDF

Similar papers

Open access Sep 2026

Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research

In this study, a comprehensive evaluation of abstractive and extractive summarization performance across three prominent large language models (LLMs): ChatGPT, DeepSeek, and Gemini is presented. A total of 8,000 cardiovascular-related research abstracts were collected from PubMed and summarized using two distinct promp...

Burcu Baştürk, Aytuğ Onan · 0 citations
Open access 2026

Knowledge Distillation for Biomedical Text Classification: A Systematic Comparative Analysis of Multiple Teacher–Student Architectures

Findings demonstrate that compact models can achieve strong biomedical classification performance through KD under compatible teacher–student pairings, while also highlighting that KD effectiveness varies substantially depending on the specific model combination.

Amine Gonca Toprak, Aytuğ Onan · 0 citations
Open access Sep 2026

A Comparative Benchmark of Biomedical Language Models for Concept Normalization from Real-World Text

A benchmark-guided, scalable framework for automated medical terminology standardization that accepts heterogeneous short medical expressions without manual input pre-processing and automatically performs text refinement, semantic retrieval and terminology mapping to standardized concepts and vocabulary codes is establ...

Anshul Verma, Abhijay, Manan Vangani et al. · 0 citations
Open access 2026

HyBioSum: A Hybrid Framework for Biomedical Long-Document Summarization With Supervised Extractive and LLM-Based Generation

A hybrid extract-then-summarize framework that first identifies salient sentences using a supervised extractive model and then generates an abstractive summary through LLM prompting is proposed, which improves efficiency by focusing the generative process on informative content while reducing the processing cost typica...

Azzedine Aftiss, Salima Lamsiyah, Christoph Schommer et al. · 0 citations
#large language models Review Open access Sep 2026

Large Language Models for Clinical Note Simplification: A Systematic Review and Experimental Evaluation of Medical Text Readability.

The findings suggest that conventional readability metrics should be extended with domain-specific measures to more accurately assess comprehensibility in medical texts and that large Language Models show strong potential to enhance the accessibility of clinical documentation for patients.

M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al. · 0 citations
2026

Assessing Transformer Models for Abstractive Summarization of Scientific Articles

The results show that BART achieves the best performance with an ROUGE-2 F1-score of 0.40664, while T5 demonstrates superior grammatical acceptability, achieving 93.36%, but BART achieves a very near performance to T5.

Emad Nabil · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.