UNMASK: Automated Detection and Mitigation of Spurious Text Correlations
August 9, 2026
UNMASK is an automated pipeline that discovers and causally verifies spurious surface patterns in text classifiers. It uses boolean expressions to identify features that boost benchmarks without providing true linguistic or causal relevance, enabling mitigation without manual annotation.
HOW THIS AFFECTS YOU
●
researcherYou can use this to identify why models fail on out-of-distribution data caused by dataset-level shortcuts.