FuzzingBrain-Bench Evaluates Open-Ended Bug Discovery via Crash Signatures
August 27, 2026
FuzzingBrain-Bench assesses LLM bug discovery by measuring the number of distinct crash signatures generated within a sanitizer-instrumented Docker environment. Unlike existing benchmarks that require specific proof-of-concept inputs for predefined vulnerabilities, this method rewards the discovery of any valid crash in open-source software.
HOW THIS AFFECTS YOU
●
researcherYou can use this to evaluate how well models actually find novel bugs rather than just replicating known vulnerabilities.