NE-BERT Improves Performance Across Nine Northeast Indian Languages
August 20, 2026
NE-BERT is a multilingual encoder trained on 8.3 million sentences to support nine underrepresented Northeast Indian languages. It achieves up to 15.97X lower perplexity than IndicBERT-V2 through aggressive upsampling and a custom SentencePiece Unigram tokenizer.
HOW THIS AFFECTS YOU
●
builderYou can utilize this model for more accurate NLP applications in linguistically diverse Indian regions.
●
researcherThis model offers a benchmark for addressing vocabulary fragmentation in extremely low-resource linguistic contexts.