New 10,000 Sentence Bangla Corpus for Function Classification
September 15, 2026
A new dataset of 10,000 manually annotated Bangla sentences across four functional categories—declarative, interrogative, imperative, and exclamatory—has been released. Benchmarking shows ensemble models like SLE and DLE outperform standard BoW and Word2Vec approaches, achieving high reliability with a Fleiss' Kappa of 0.82.
HOW THIS AFFECTS YOU
●
builderThis provides a foundation for building Bangla-specific dialogue or TTS systems.
●
researcherYou can use this balanced corpus to benchmark multilingual NLP models.