Contextual Models Replace Lexicons for Multilingual Grievance Labeling
July 24, 2026
Replacing word-level lexicons with context-reading models corrects a massive bias in grievance detection where AUROC scores were artificially inflated by lexicon construction. The new approach uses a five-language, 2,000-item evaluation pool to resolve negation and quotation in threat assessment.
HOW THIS AFFECTS YOU
●
researcherYou can avoid the circular evaluation errors common in lexicon-based threat detection.
●
policyYou can more accurately monitor online threats by using models that understand linguistic context rather than simple keyword matching.