The APEX-VW (Automatic Post-Editing eXperiments on Virtual Wards) Corpus is presented, a new open English-Spanish (EN-ES) dataset built from recent NHS virtual-ward documents and professional PE in Trados Studio, with controlled MT, terminology, and quality assurance settings.
Abstract
Post-Editing (PE) of Machine Translation (MT) output often involves repeating the same lexical and terminological corrections across many segments, especially in specialised and highly repetitive documents. Despite substantial work on Automatic Post-Editing (APE), most available corpora operate at the sentence level, others are synthetic, and overall not designed to study how corrections propagate in realistic Computer-Assisted Translation (CAT) workflows. This paper presents the APEX-VW (Automatic Post-Editing eXperiments on Virtual Wards) Corpus, a new open English-Spanish (EN-ES) dataset built from recent NHS virtual-ward documents and professional PE in Trados Studio, with controlled MT, terminology, and quality assurance settings. The corpus contains seven document-coherent source texts totalling 42k words, translated with four MT systems representing different paradigms and then post-edited by professional translators. Unlike prior resources such as WMT APE corpora, eSCAPE, MLQE-PE, or LangMark, the dataset preserves document order and CAT-tool context, making it suitable for research on terminology normalisation, correction propagation, and human-in-the-loop translation support. The paper describes the corpus design, data preparation, PE setup, and initial corpus statistics, and positions the resource as a benchmark for document-level APE and propagation-aware assistive tools.
En-ViMedNER is presented, the first English-Vietnamese parallel biomedical NER corpus annotated with UMLS semantic types, which are language-neutral codes providing a shared cross-lingual label space and ensuring direct comparability with existing UMLS-based resources.
Nhu Vo, P. Nguyen, Nu-Uyen-Phuong Le et al.· 0 citations
Background: To determine whether cross-lingual clinical annotation projection can be formulated as a text-preserving, document-level generative task that produces verifiable character-level annotations for multilingual clinical corpus construction, and to characterize its robustness and computational trade-offs relativ...
A. Rey-Blanes, F. J. Moreno-Barea, F. J. Veredas· 0 citations
The University of Amsterdam’s participation in the TREC 2024 Plain Language Adaptation of Biomedical Abstracts (PLABA) Track is documented and the effectiveness of text simplification models trained on aligned pairs of sentences in biomedical abstracts and plain language summaries is investigated.
Jan Bakker, Taiki Papandreou-Lazos, Jaap Kamps· 0 citations
A benchmark-guided, scalable framework for automated medical terminology standardization that accepts heterogeneous short medical expressions without manual input pre-processing and automatically performs text refinement, semantic retrieval and terminology mapping to standardized concepts and vocabulary codes is establ...
Anshul Verma, Abhijay, Manan Vangani et al.· bioRxiv· 0 citations
The proposed Semantic-Syntax Prealignment (SSPA), an innovative corpus generation framework, exhibits remarkable cross-domain adaptability and stylistic consistency, offering a robust, versatile solution for low-resource Tibetan professional domain MT.