Vero Benchmark for Repository-Level Verified Code Generation
August 14, 2026
Vero is a new benchmark for evaluating AI agents' ability to perform joint implementation and formal proof synthesis across multi-module repositories. It includes 43 real-world instances in Python, Dafny, Verus, and Coq, covering distributed systems and cryptography.
HOW THIS AFFECTS YOU
●
builderThis provides a way to measure the reliability of agents writing mission-critical, formally verified software.
●
researcherYou can use Vero to test if agents can maintain logical consistency across complex, multi-file codebases.