Skip to content
Book Open access

BioCLEAR Benchmark for Biomedical Text Simplification

Jul 2026 · Annual International ACM SIGIR Conference on Research and Development in Information Retrieval · pp. 3101-3108 · 0 citations · 25 references
Computer Science

TL;DR

BioCLEAR is constructed, a benchmark for Biomedical Comprehensible Lay Explanations of Academic Research used in the CLEF 2025 SimpleText track, comprising a large set of aligned abstracts and plain-language summaries from Cochrane systematic reviews, which provides a new resource for scientific document-level text simplification.

Abstract

Access to objective, reliable scientific information is crucial in a world of misinformation and disinformation, yet the general public often avoids scientific literature due to its perceived complexity. Modern generative information access models hold the promise of removing some of these barriers by developing appropriate approaches to scientific text simplification. We constructed BioCLEAR, a benchmark for Biomedical Comprehensible Lay Explanations of Academic Research used in the CLEF 2025 SimpleText track, comprising a large set of aligned abstracts and plain-language summaries from Cochrane systematic reviews. First, it provides a new resource for scientific document-level text simplification, including discourse-level aspects that are absent from traditional sentence-level data. Second, we include existing resources created through human sentence-level simplification to run experiments across different datasets. Third, it includes a large set of sentence-level and document-level submissions to the CLEF track, providing a sample of current models and enabling detailed analysis of their effectiveness and remaining issues.

Read PDF

Similar papers

#large language models Review Open access Sep 2026

Large Language Models for Clinical Note Simplification: A Systematic Review and Experimental Evaluation of Medical Text Readability.

The findings suggest that conventional readability metrics should be extended with domain-specific measures to more accurately assess comprehensibility in medical texts and that large Language Models show strong potential to enhance the accessibility of clinical documentation for patients.

M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al. · 0 citations
Open access 2026

Knowledge Distillation for Biomedical Text Classification: A Systematic Comparative Analysis of Multiple Teacher–Student Architectures

Findings demonstrate that compact models can achieve strong biomedical classification performance through KD under compatible teacher–student pairings, while also highlighting that KD effectiveness varies substantially depending on the specific model combination.

Amine Gonca Toprak, Aytuğ Onan · 0 citations

Data Gatherer Revisited: Scalable Dataset Reference Extraction from Biomedical Literature

This work improves evidence retrieval by combining semantic ranking and repository-specific identifier pattern matching and achieves recall close to that of proprietary models while reducing estimated inference cost by 330× to 1,300×, although with lower precision.

Pietro Marini, M. Peters, Nicole Contaxis et al. · 0 citations
Open access Sep 2026

A Comparative Benchmark of Biomedical Language Models for Concept Normalization from Real-World Text

A benchmark-guided, scalable framework for automated medical terminology standardization that accepts heterogeneous short medical expressions without manual input pre-processing and automatically performs text refinement, semantic retrieval and terminology mapping to standardized concepts and vocabulary codes is establ...

Anshul Verma, Abhijay, Manan Vangani et al. · 0 citations

Optimizing large language model prompts for biomedical knowledge discovery

This work presents a scalable, reproducible framework for evaluating, optimizing, and interpreting LLMs for biomedical knowledge extraction, with a focus on gene–gene regulatory relation prediction, pathway component recognition, multimodal pathway figure understanding, and automated prompt optimization.

Muhammad Azam · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.