Skip to content
Preprint

APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain

Aug 2026 · 0 citations · 23 references
Computer Science

TL;DR

The APEX-VW (Automatic Post-Editing eXperiments on Virtual Wards) Corpus is presented, a new open English-Spanish (EN-ES) dataset built from recent NHS virtual-ward documents and professional PE in Trados Studio, with controlled MT, terminology, and quality assurance settings.

Abstract

Post-Editing (PE) of Machine Translation (MT) output often involves repeating the same lexical and terminological corrections across many segments, especially in specialised and highly repetitive documents. Despite substantial work on Automatic Post-Editing (APE), most available corpora operate at the sentence level, others are synthetic, and overall not designed to study how corrections propagate in realistic Computer-Assisted Translation (CAT) workflows. This paper presents the APEX-VW (Automatic Post-Editing eXperiments on Virtual Wards) Corpus, a new open English-Spanish (EN-ES) dataset built from recent NHS virtual-ward documents and professional PE in Trados Studio, with controlled MT, terminology, and quality assurance settings. The corpus contains seven document-coherent source texts totalling 42k words, translated with four MT systems representing different paradigms and then post-edited by professional translators. Unlike prior resources such as WMT APE corpora, eSCAPE, MLQE-PE, or LangMark, the dataset preserves document order and CAT-tool context, making it suitable for research on terminology normalisation, correction propagation, and human-in-the-loop translation support. The paper describes the corpus design, data preparation, PE setup, and initial corpus statistics, and positions the resource as a benchmark for document-level APE and propagation-aware assistive tools.

View source

Similar papers

#natural language process... Preprint Aug 2026

En-ViMedNER: An English-Vietnamese Parallel Biomedical Corpus with UMLS Semantic Type Annotations

En-ViMedNER is presented, the first English-Vietnamese parallel biomedical NER corpus annotated with UMLS semantic types, which are language-neutral codes providing a shared cross-lingual label space and ensuring direct comparability with existing UMLS-based resources.

Nhu Vo, P. Nguyen, Nu-Uyen-Phuong Le et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language Study

Background: To determine whether cross-lingual clinical annotation projection can be formulated as a text-preserving, document-level generative task that produces verifiable character-level annotations for multilingual clinical corpus construction, and to characterize its robustness and computational trade-offs relativ...

A. Rey-Blanes, F. J. Moreno-Barea, F. J. Veredas · 0 citations

UvA-DARE (Digital Academic Repository) Biomedical Text Simplification Models Trained on Aligned Abstracts and Lay Summaries

The University of Amsterdam’s participation in the TREC 2024 Plain Language Adaptation of Biomedical Abstracts (PLABA) Track is documented and the effectiveness of text simplification models trained on aligned pairs of sentences in biomedical abstracts and plain language summaries is investigated.

Jan Bakker, Taiki Papandreou-Lazos, Jaap Kamps · 0 citations
Open access Sep 2026

A Comparative Benchmark of Biomedical Language Models for Concept Normalization from Real-World Text

A benchmark-guided, scalable framework for automated medical terminology standardization that accepts heterogeneous short medical expressions without manual input pre-processing and automatically performs text refinement, semantic retrieval and terminology mapping to standardized concepts and vocabulary codes is establ...

Anshul Verma, Abhijay, Manan Vangani et al. · 0 citations
Open access Sep 2026

SSPA: Enhancing Pseudo-Corpus Quality on Tibetan Machine Translation via Semantic-Syntax Prealignment

The proposed Semantic-Syntax Prealignment (SSPA), an innovative corpus generation framework, exhibits remarkable cross-domain adaptability and stylistic consistency, offering a robust, versatile solution for low-resource Tibetan professional domain MT.

Yi-Dong Sun, Dong-Xu Liu, Jia-Lei Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.