Skip to content
Preprint

PlainMedScale: A Corpus of Multi-Level Simplified Medical Texts in German and English

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

PlainMedScale, a topic-aligned medical corpus spanning four levels of comprehensibility in German and English, drawn from MSD, Gesund.Bund, Apotheken Umschau Einfache Sprache, and the NHS is introduced.

Abstract

We introduce PlainMedScale, a topic-aligned medical corpus spanning four levels of comprehensibility in German and English, drawn from MSD (professional and consumer), Gesund.Bund, Apotheken Umschau Einfache Sprache, and the NHS. The four tiers correspond to distinct communicative functions --- reference, explanation, decision support, and access --- and move beyond the binary expert--lay contrast of prior corpora. In two pilot studies enabled by the alignments, we show that many readability metrics established on two registers fail to generalize across the full gradient, and that a SOTA open-weight LLM prompted for Plain Language still partially preserves the difficulty of its input. Code (https://github.com/GS-Uni-Heidelberg/PlainMedScale) and data (https://doi.org/10.5281/zenodo.21728290) are made available.

View source

Similar papers

Open access Aug 2026

Benchmark of small language models for plain-language simplification in Spanish clinical texts

This work presents the first benchmark of small language models for Spanish clinical plain-language adaptation and introduces MEDICLARO, a corpus specifically designed for this task, and establishes a solid foundation for integrating small language models into Spanish clinical workflows.

P. Martínez, Jesús M. Sánchez-Gómez, Lourdes Moreno · 1 citation
Review Aug 2026

HealMed: Multilingual Evaluation of Large Language Models in Medicine

On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models, whereas many open-source and medically specialized models showed larger and less consistent gaps.

Yingjian Chen, Fan Gao, Sherry T. Tong et al. · 0 citations
Preprint Aug 2026

APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain

The APEX-VW (Automatic Post-Editing eXperiments on Virtual Wards) Corpus is presented, a new open English-Spanish (EN-ES) dataset built from recent NHS virtual-ward documents and professional PE in Trados Studio, with controlled MT, terminology, and quality assurance settings.

Marie Escribe, Tharindu Ranasinghe, A. Haddad et al. · 0 citations
#large language models Review Open access Sep 2026

Large Language Models for Clinical Note Simplification: A Systematic Review and Experimental Evaluation of Medical Text Readability.

The findings suggest that conventional readability metrics should be extended with domain-specific measures to more accurately assess comprehensibility in medical texts and that large Language Models show strong potential to enhance the accessibility of clinical documentation for patients.

M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al. · 0 citations
#natural language process... Preprint Aug 2026

En-ViMedNER: An English-Vietnamese Parallel Biomedical Corpus with UMLS Semantic Type Annotations

En-ViMedNER is presented, the first English-Vietnamese parallel biomedical NER corpus annotated with UMLS semantic types, which are language-neutral codes providing a shared cross-lingual label space and ensuring direct comparability with existing UMLS-based resources.

Nhu Vo, P. Nguyen, Nu-Uyen-Phuong Le et al. · 0 citations
Open access Aug 2026

When Artificial Intelligence Speaks For The Obstetrician: Multilingual Accuracy On Real Patient Questions

This study aimed to compare the response accuracy and reference quality of three free LLMs in Turkish and English using common pregnancy-related questions and found that Google Gemini and DeepSeek provided more accurate responses in English than in Turkish.

A. Tığlı, Y. Baykuş, R. Deniz et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.