KaliBench provides a fine-grained dataset of 8,504 natural-language-to-CLI pairs across 1,642 tools on Kali Linux. It measures the ability of LLMs to generate syntactically correct, executable commands for complex security workflows.
HOW THIS AFFECTS YOU
●
builderYou can use this to benchmark the reliability of LLM-based agents performing cybersecurity operations.
●
researcherThis allows for more precise, verifiable evaluation of tool-use capabilities in security contexts.