Tokenization and Encoding Biases in Crosslingual LLM Evaluation
August 27, 2026
An empirical study reveals that widely used normalized metrics in crosslingual evaluation introduce significant biases due to tokenization, encoding, and orthographic differences. The research suggests sentence-level negative log-likelihood over semantically equivalent pairs as a more robust alternative.
HOW THIS AFFECTS YOU
●
researcherYou should avoid standard normalized metrics when comparing model performance across different languages.