Cultivar Benchmark Probes Translation Contamination and Localisation
August 9, 2026
Cultivar is a source-contrastive translation benchmark that uses a localized subset of FLORES to evaluate linguistic robustness. Testing 32 open-weight models revealed that MT-specialized models are less robust and frequently exhibit US-centric bias regardless of the target language.
HOW THIS AFFECTS YOU
●
researcherYou can use this to detect if your multilingual models are overfit to standard English-centric datasets.
●
policyThis highlights potential cultural biases in AI translation that could impact regional communication standards.