Waldo Benchmark Identifies Multilingual Query-Language Bias in LLMs
October 2, 2026
The Waldo benchmark comprises 12,000 Wikipedia-based QA pairs designed to identify knowledge disparities across languages. It measures how LLMs exhibit query-language preference, which can lead to inconsistent or incomplete information retrieval when facts are only available in certain languages.
HOW THIS AFFECTS YOU
●
researcherYou can use Waldo to evaluate how models handle information asymmetry in multilingual settings.
●
policyThis highlights potential biases in how different linguistic populations receive factual information from AI.