●builderThis provides a method to defend your chat agents against multi-turn jailbreaking attempts.
●researcherYou can apply token-level weighting based on refusal-attributable advantage to improve safety training.
●policyThis advances technical safeguards against complex, multi-stage model misuse.