HyBioSum: A Hybrid Framework for Biomedical Long-Document Summarization With Supervised Extractive and LLM-Based Generation
Abstract
Biomedical text summarization (BTS) aims to automatically condense single or multiple biomedical documents into concise summaries while preserving essential information. However, biomedical texts are often lengthy, structurally complex, and rich in domain-specific terminology, making automated summarization particularly challenging. Although recent Large Language Models (LLMs) have demonstrated strong performance across various NLP tasks, their direct application to biomedical documents remains difficult due to computational constraints and the risk of losing clinically important information in long inputs. To address these limitations, we propose a hybrid extract-then-summarize framework that first identifies salient sentences using a supervised extractive model and then generates an abstractive summary through LLM prompting. This strategy improves efficiency by focusing the generative process on informative content while reducing the processing cost typically associated with long biomedical documents. We evaluate the proposed approach on benchmark datasets, including PubMed and CORD-19, using ROUGE, METEOR, and BERTScore metrics. Additionally, we conduct a human evaluation assessing relevance, conciseness, informativeness, and readability. Experimental results show that our method achieves competitive performance compared to recent state-of-the-art models. Ablation studies further confirm the benefits of integrating supervised extraction with LLM-based generation within a unified hybrid framework.