PlainMedScale, a topic-aligned medical corpus spanning four levels of comprehensibility in German and English, drawn from MSD, Gesund.Bund, Apotheken Umschau Einfache Sprache, and the NHS is introduced.
Abstract
We introduce PlainMedScale, a topic-aligned medical corpus spanning four levels of comprehensibility in German and English, drawn from MSD (professional and consumer), Gesund.Bund, Apotheken Umschau Einfache Sprache, and the NHS. The four tiers correspond to distinct communicative functions --- reference, explanation, decision support, and access --- and move beyond the binary expert--lay contrast of prior corpora. In two pilot studies enabled by the alignments, we show that many readability metrics established on two registers fail to generalize across the full gradient, and that a SOTA open-weight LLM prompted for Plain Language still partially preserves the difficulty of its input. Code (https://github.com/GS-Uni-Heidelberg/PlainMedScale) and data (https://doi.org/10.5281/zenodo.21728290) are made available.
This work presents the first benchmark of small language models for Spanish clinical plain-language adaptation and introduces MEDICLARO, a corpus specifically designed for this task, and establishes a solid foundation for integrating small language models into Spanish clinical workflows.
P. Martínez, Jesús M. Sánchez-Gómez, Lourdes Moreno· Scientific Reports· 1 citation
On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models, whereas many open-source and medically specialized models showed larger and less consistent gaps.
Yingjian Chen, Fan Gao, Sherry T. Tong et al.· 0 citations
The APEX-VW (Automatic Post-Editing eXperiments on Virtual Wards) Corpus is presented, a new open English-Spanish (EN-ES) dataset built from recent NHS virtual-ward documents and professional PE in Trados Studio, with controlled MT, terminology, and quality assurance settings.
Marie Escribe, Tharindu Ranasinghe, A. Haddad et al.· 0 citations
The findings suggest that conventional readability metrics should be extended with domain-specific measures to more accurately assess comprehensibility in medical texts and that large Language Models show strong potential to enhance the accessibility of clinical documentation for patients.
M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al.· Studies in Health Technology...· 0 citations
En-ViMedNER is presented, the first English-Vietnamese parallel biomedical NER corpus annotated with UMLS semantic types, which are language-neutral codes providing a shared cross-lingual label space and ensuring direct comparability with existing UMLS-based resources.
Nhu Vo, P. Nguyen, Nu-Uyen-Phuong Le et al.· 0 citations
This study aimed to compare the response accuracy and reference quality of three free LLMs in Turkish and English using common pregnancy-related questions and found that Google Gemini and DeepSeek provided more accurate responses in English than in Turkish.
A. Tığlı, Y. Baykuş, R. Deniz et al.· Hippocrates Medical Journal· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.