C-Guard: Data-Efficient RL Alignment for Conflicting Objectives
August 4, 2026
C-Guard uses a constitution-grid instrument and a per-cell learnability score (C-LIM) to optimize RL alignment. The method identifies and prunes low-impact data, improving learning impact from 0.733 to 0.80 and mitigating the trade-off between over-refusal and under-refusal.
HOW THIS AFFECTS YOU
●
builderThis method offers a more data-efficient way to tune safety guardrails without increasing over-refusal.
●
researcherYou can use C-LIM to identify dead-weight data regions before committing training budgets.