Training-time explainability improves multilingual hate speech detection
August 28, 2026
A new framework aligns model reasoning with human-annotated rationales during training to improve hate speech detection in English and Hinglish. Using gradient- and attention-based regularization, the method captures culturally specific cues for implicit hate speech.
HOW THIS AFFECTS YOU
●
researcherYou can use training-time regularization to improve both classification F-scores and explanation faithfulness.
●
policyThis approach offers more interpretable moderation tools for culturally sensitive content.