SWORD Benchmark Reveals Cross-Lingual Factual Inconsistencies in LLMs
September 10, 2026
The SWORD benchmark uses Wikidata-based object-relation distortions to show that LLMs rely on distributional familiarity rather than factual truth. Findings indicate models are more likely to accept semantically plausible errors than random ones across eight languages.
HOW THIS AFFECTS YOU
●
researcherThis highlights a critical flaw where models prioritize linguistic plausibility over factual accuracy.
●
policyYou should be aware that multilingual models may exhibit unpredictable factual failures in cross-lingual contexts.