MaliciousSkillBench consolidates 13 sources into a benchmark of 9,740 skills, containing 7,505 malicious and 2,235 benign identities. The dataset maps 4,588 operational structural families across 11 attack categories to detect malicious instruction packages and scripts in agentic workflows.
HOW THIS AFFECTS YOU
●
builderYou can use this to test the security of your agent's skill-distribution channels.
●
policyThis provides a standardized framework for measuring risks in agentic autonomy.