The Arabic Safety Index (ASAS) is a human-curated benchmark comprising 801 prompts across eight safety categories for redteaming Arabic LLMs. Evaluations of models like GPT-4o and Claude 3.7 Sonnet show they fail to defend against 50% of unsafe prompts in high-harm categories.
HOW THIS AFFECTS YOU
●
researcherUse ASAS to evaluate cultural alignment and safety gaps in Arabic-capable models.
●
policyThe findings highlight significant safety vulnerabilities in regional LLM deployments.