DECO Framework Reveals Hidden Failures in LLM Content Moderation
September 4, 2026
The DECO framework factorizes content moderation into independent criteria to reveal that high benchmark scores often mask failures in specific task dimensions. Testing across four LLMs shows models struggle when decisions depend on specific aspects rather than general harmfulness.
HOW THIS AFFECTS YOU
●
builderYou should implement criterion-level evaluations to ensure your moderation layer is reliable across all safety dimensions.
●
policyBe cautious of aggregate moderation scores, as they may hide critical failures in specific safety dimensions.