mDeBERTa-v3 Reaches 0.841 F1 for Sinhala-to-English Hallucination Detection
October 9, 2026
A new reference-free framework detects semantic hallucinations in low-resource Sinhala-to-English translation using a 45,000-sample synthetic dataset. The fine-tuned mDeBERTa-v3 model achieves a token-level F1 score of 0.841 by utilizing linguistically motivated corruption strategies.
HOW THIS AFFECTS YOU
●
builderUse this framework to improve translation reliability in low-resource language pipelines.
●
researcherThe five-stage corruption strategy provides a new method for generating synthetic hallucination data.