CleanSlate Benchmark for Machine Unlearning Forget Set Curation
August 18, 2026
CleanSlate addresses the challenge of mapping suppression requests to specific training data in massive corpora. The benchmark uses model-specific extraction and content-grounded QA to show that standard lexical curators fail to achieve effective verbatim output suppression.
HOW THIS AFFECTS YOU
●
builderYou cannot rely on simple substring matching to satisfy data deletion requests.
●
researcherThis highlights the gap between known forget sets and real-world data curation.