Disjoint Token Spaces Impede Cross-Lingual Knowledge Transfer in LLMs
September 18, 2026
Research using 360M and 7B parameter models reveals that disjoint token spaces cause knowledge compartmentalization during pretraining. Even when training on identical text, different token mappings prevent models from sharing linguistic knowledge across languages.
HOW THIS AFFECTS YOU
●
researcherThis suggests that tokenization strategies are a fundamental bottleneck for multilingual generalization, independent of training data quality.