Terminal-Bench-Science Evaluates AI Agents on Expert Scientific Workflows
August 27, 2026
Terminal-Bench-Science 0.1 introduces a benchmark of 70 expert-curated tasks across life, physical, and engineering sciences. Claude Opus 5 currently leads with a 30% resolution rate on these complex scientific research workflows.
HOW THIS AFFECTS YOU
●
builderThis provides a signal for the performance of coding and reasoning agents in domain-specific environments.
●
researcherYou can now benchmark agentic capabilities against specialized scientific tasks rather than general reasoning.