Automated NER Annotation Correction for Low-Resource Languages
September 17, 2026
This framework uses a frequency-based iterative approach with self-training and a dual-threshold mechanism to clean noisy NER datasets. The method improves performance in low-resource settings and explores using generative LLMs to perform NER tasks directly.
HOW THIS AFFECTS YOU
●
builderYou can improve the quality of your NLP training data in low-resource domains using automated iterative cleaning.
●
researcherThe dual-threshold mechanism offers a new way to handle label noise in NER datasets.