This new corpus consists of 42k words from NHS virtual-ward documents to study document-level machine translation post-editing. It provides a realistic dataset for evaluating how professional translators correct terminological and lexical errors across coherent document segments.
HOW THIS AFFECTS YOU
●
builderYou can use this dataset to fine-tune translation models specifically for medical document workflows.
●
healthThis improves the reliability of automated translation in clinical documentation settings.