TransClean benchmark for LLM translation noise detection
September 11, 2026
TransClean establishes a benchmark of 9,900 noisy and clean translation pairs to address LLM-generated translation noise like language labels and explanations. The study analyzes 790,000 outputs from 12 LLMs to evaluate span-based and LLM-based extraction methods.
HOW THIS AFFECTS YOU
●
builderUse this benchmark to improve the reliability of LLM-based translation pipelines by filtering out unwanted metadata.
●
researcherThis provides a systematic framework for studying and quantifying translation noise in large-scale model outputs.