BioStudyBench Evaluates Agents on Post-Cutoff Biomedical Data
October 7, 2026
BioStudyBench consists of 25 long-horizon tasks requiring agents to find and analyze public data from studies published after their knowledge cutoffs. Testing across eight models showed that providing access to tools and data increased pass rates by 47 percentage points.
HOW THIS AFFECTS YOU
●
researcherYou can use this benchmark to test if agents are truly performing analysis rather than retrieving training data.
●
healthYou can evaluate the reliability of autonomous agents in performing actual biomedical research tasks.