This research introduces Operative Understanding Rate (UR) to supplement Attack Success Rate (ASR) in safety benchmarks. Measuring UR allows evaluators to distinguish between models that correctly refuse harmful tasks and those that simply fail to engage with the prompt intent.
HOW THIS AFFECTS YOU
●
researcherYou can use UR to prevent misleadingly low ASR results from masking poor model reasoning.
●
policyThis offers a more nuanced metric for auditing model safety and alignment compliance.