HarmProfile provides a content-centric benchmark that treats model misbehavior as a distribution of harm categories. It allows researchers to define a specific risk profile based on the content and severity of safety failures.
HOW THIS AFFECTS YOU
●
researcherYou can now characterize model risk through content variation rather than just attack outcomes.
●
policyThis provides a more granular framework for evaluating frontier model safety.