Multilingual Model Training Decreases Lexical Normalization Accuracy
September 2, 2026
Training a single fixed-capacity character-level model on up to twelve languages causes a significant performance drop in lexical normalization. Accuracy decreases by approximately 40% as more languages are added, particularly when total training data remains constant.
HOW THIS AFFECTS YOU
●
researcherYou should limit the number of languages in joint training to 1–4 to avoid capacity competition and accuracy loss.