LLM Performance in MQM and ESA Machine Translation Quality Annotation
October 9, 2026
Evaluation of LLMs across 70 language pairs shows that model agreement with human annotators can exceed inter-human agreement for MQM and ESA schemes. The study uses WMT23 and WMT25 datasets to assess reliability across different domains and annotation granularities.
HOW THIS AFFECTS YOU
●
builderYou can leverage LLMs for cost-effective MT quality monitoring if you account for specific language pair variances.
●
researcherYou can use these findings to calibrate LLM-based evaluation metrics against human benchmarks.