LENS protocol evaluates narrative unlearning in LLMs
July 28, 2026
The LENS protocol introduces a multi-level evaluation for machine unlearning, measuring how well models suppress specific disinformation narratives without total model collapse. It uses a Suppression-Collapse Efficiency (SCE) score to balance target suppression against general output quality across 12B-parameter models.
HOW THIS AFFECTS YOU
●
researcherThe SCE score provides a more nuanced metric for evaluating the effectiveness of unlearning algorithms.
●
policyThis helps quantify the risks and efficacy of attempting to programmatically remove harmful disinformation from models.