Jul 2026· Annual International ACM SIGIR Conference on Research and Development in Information Retrieval· pp. 3101-3108· 0 citations· 25 references
Computer Science
TL;DR
BioCLEAR is constructed, a benchmark for Biomedical Comprehensible Lay Explanations of Academic Research used in the CLEF 2025 SimpleText track, comprising a large set of aligned abstracts and plain-language summaries from Cochrane systematic reviews, which provides a new resource for scientific document-level text simplification.
Abstract
Access to objective, reliable scientific information is crucial in a world of misinformation and disinformation, yet the general public often avoids scientific literature due to its perceived complexity. Modern generative information access models hold the promise of removing some of these barriers by developing appropriate approaches to scientific text simplification. We constructed BioCLEAR, a benchmark for Biomedical Comprehensible Lay Explanations of Academic Research used in the CLEF 2025 SimpleText track, comprising a large set of aligned abstracts and plain-language summaries from Cochrane systematic reviews. First, it provides a new resource for scientific document-level text simplification, including discourse-level aspects that are absent from traditional sentence-level data. Second, we include existing resources created through human sentence-level simplification to run experiments across different datasets. Third, it includes a large set of sentence-level and document-level submissions to the CLEF track, providing a sample of current models and enabling detailed analysis of their effectiveness and remaining issues.
The findings suggest that conventional readability metrics should be extended with domain-specific measures to more accurately assess comprehensibility in medical texts and that large Language Models show strong potential to enhance the accessibility of clinical documentation for patients.
M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al.· Studies in Health Technology...· 0 citations
These findings provide a systematic reference for selecting biomedical summarization tools and highlight that broad pretraining outperforms narrow domain adaptation.
Fabio Baumgärtel, Enrico Bono, Lucas Fillinger et al.· iScience· 0 citations
Findings demonstrate that compact models can achieve strong biomedical classification performance through KD under compatible teacher–student pairings, while also highlighting that KD effectiveness varies substantially depending on the specific model combination.
This work improves evidence retrieval by combining semantic ranking and repository-specific identifier pattern matching and achieves recall close to that of proprietary models while reducing estimated inference cost by 330× to 1,300×, although with lower precision.
Pietro Marini, M. Peters, Nicole Contaxis et al.· 0 citations
A benchmark-guided, scalable framework for automated medical terminology standardization that accepts heterogeneous short medical expressions without manual input pre-processing and automatically performs text refinement, semantic retrieval and terminology mapping to standardized concepts and vocabulary codes is establ...
Anshul Verma, Abhijay, Manan Vangani et al.· bioRxiv· 0 citations
This work presents a scalable, reproducible framework for evaluating, optimizing, and interpreting LLMs for biomedical knowledge extraction, with a focus on gene–gene regulatory relation prediction, pathway component recognition, multimodal pathway figure understanding, and automated prompt optimization.
Muhammad Azam· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.