SWE-rebench v2 Leaderboard for Software Engineering Agents
July 31, 2026
The SWE-rebench v2 benchmark evaluates 13 models and 4 agents on coding tasks across Go, Java, Python, Rust, and TS. Anthropic Fable leads with a 64.5% resolved rate, while OpenAI GPT-5.6 Sol shows high efficiency at $0.85 per problem.
HOW THIS AFFECTS YOU
●
builderYou can benchmark your autonomous coding agents against industry leaders using multi-language repository data.
●
researcherThis provides a standardized evaluation framework for agentic reasoning in complex software engineering workflows.