SWE-bench Science Benchmark Shows Sub-50% Pass Rate for Agents
August 21, 2026
SWE-bench Science evaluates scientific software engineering across 119 tasks in 98 repositories. Even top-tier agents like Claude Code with Opus-5 achieve a pass@1 rate of less than 50%.
HOW THIS AFFECTS YOU
●
builderExpect limited reliability when deploying autonomous agents for complex scientific software tasks.
●
researcherThis benchmark identifies a significant performance gap in agentic reasoning within scientific domains.