SWE-Prometheus moves beyond functional patch testing to evaluate how coding agents handle repository engineering governance, including risk identification and intervention prioritization. Testing across 60 repositories showed model Normalized Governance Improvement scores ranging from 0.0568 to 0.5760.
HOW THIS AFFECTS YOU
●
builderYou can use this benchmark to assess how well your coding agents manage technical debt and risk.
●
researcherThis provides a more holistic evaluation metric for agentic software engineering than simple pass/fail patches.