These findings provide a systematic reference for selecting biomedical summarization tools and highlight that broad pretraining outperforms narrow domain adaptation.
Abstract
Summary The rapid expansion of biomedical literature demands automated summarization tools that reliably condense research articles into concise, accurate summaries. We benchmarked 62 summarization methods, ranging from frequency-based and TextRank extractors to encoder-decoder models (EDMs) and large language models (LLMs), on 1,000 biomedical abstracts from 20 journals across ScienceDirect and Cell Press, using author-written highlights as reference summaries. Models were evaluated with a composite suite of lexical, semantic, and factual metrics, including ROUGE, BLEU, METEOR, embedding-based similarity, and factuality scores. General-purpose models (e.g., Mistral, GPT, and Llama) achieved the highest overall performance across lexical and semantic dimensions, outperforming reasoning-oriented (e.g., DeepSeek and Magistral) and domain-specific (e.g., BioGPT and BioMistral) models. Notably, medium-sized models outperformed large-scale models, suggesting an optimal balance between model capacity and efficiency, while classical extractive methods lagged behind neural approaches. These findings provide a systematic reference for selecting biomedical summarization tools and highlight that broad pretraining outperforms narrow domain adaptation.
In this study, a comprehensive evaluation of abstractive and extractive summarization performance across three prominent large language models (LLMs): ChatGPT, DeepSeek, and Gemini is presented. A total of 8,000 cardiovascular-related research abstracts were collected from PubMed and summarized using two distinct promp...
Burcu Baştürk, Aytuğ Onan· Sakarya University Journal o...· 0 citations
Findings demonstrate that compact models can achieve strong biomedical classification performance through KD under compatible teacher–student pairings, while also highlighting that KD effectiveness varies substantially depending on the specific model combination.
A benchmark-guided, scalable framework for automated medical terminology standardization that accepts heterogeneous short medical expressions without manual input pre-processing and automatically performs text refinement, semantic retrieval and terminology mapping to standardized concepts and vocabulary codes is establ...
Anshul Verma, Abhijay, Manan Vangani et al.· bioRxiv· 0 citations
A hybrid extract-then-summarize framework that first identifies salient sentences using a supervised extractive model and then generates an abstractive summary through LLM prompting is proposed, which improves efficiency by focusing the generative process on informative content while reducing the processing cost typica...
Azzedine Aftiss, Salima Lamsiyah, Christoph Schommer et al.· IEEE Access· 0 citations
The findings suggest that conventional readability metrics should be extended with domain-specific measures to more accurately assess comprehensibility in medical texts and that large Language Models show strong potential to enhance the accessibility of clinical documentation for patients.
M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al.· Studies in Health Technology...· 0 citations
The results show that BART achieves the best performance with an ROUGE-2 F1-score of 0.40664, while T5 demonstrates superior grammatical acceptability, achieving 93.36%, but BART achieves a very near performance to T5.
Emad Nabil· Islamic University Journal o...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.