The ASRD dataset tests open-weight models against non-canonical inputs like leetspeak, emojis, and encoded strings. Evaluations show that while emojis cause minimal comprehension failure, models still exhibit significant harmful compliance when faced with varied surface-form transformations.
HOW THIS AFFECTS YOU
●
researcherYou can use this dataset to evaluate how robust model safety guardrails are against adversarial formatting.
●
policyThis highlights a need for safety benchmarks that account for non-standard character encoding and obfuscation.