Code-Switching Curricula Improve Cross-Lingual Alignment in Small Language Models
September 28, 2026
Training small decoder-only transformers on LLM-generated code-switched data induces better cross-lingual alignment, particularly across different scripts. A curriculum progressing from word-level to sentence-level code-switching allows models to outperform monolingual training baselines on the BabyLM evaluation suite.
HOW THIS AFFECTS YOU
●
builderThis technique can improve the multilingual capabilities of small, edge-deployable models.
●
researcherYou can apply these curricula to improve representation alignment in multilingual pretraining.