CLEAR uses a lightweight hidden-state gate to control a safety low-rank adapter, allowing for conditional safety activation. This method improves robustness on HarmBench while minimizing the utility loss typically seen in standard SFT or LoRA-based safety tuning.
HOW THIS AFFECTS YOU
●
builderYou can implement more surgical safety guardrails that don't degrade model performance on benign prompts.
●
researcherYou can apply this conditional gating mechanism to decouple safety alignment from general utility.