●researcherYou can use these distinct detection methods to separate user-side toxicity from model-directed adversarial pressure.
●policyThis highlights why standard toxicity filters may fail to catch coercion and harassment targeting the model itself.