IndicTriMix for Token-Level Code-Mixed Language Identification
September 11, 2026
IndicTriMix introduces a benchmark and fine-tuned MuRIL and XLM-RoBERTa models for token-level language identification in tri-language code-mixed text. The approach treats code-mixing as a sequence labeling problem specifically optimized for Hindi, Gujarati, and Bengali.
HOW THIS AFFECTS YOU
●
builderYou can use these models and benchmarks to better process multilingual social media text in Indian markets.
●
researcherThis work provides a new methodology for generating code-mixed training data using parallel sentences.