BenchBench-Protocol Evaluates LLM Reasoning for Wet-Lab Protocol Modification
August 26, 2026
This benchmark uses 149 tasks derived from real-world scientific protocol modifications across nine biology domains. It tests an LLM's ability to adapt published protocols by accounting for downstream experimental dependencies and prior procedural choices.
HOW THIS AFFECTS YOU
●
researcherYou can use this to evaluate specialized life-science models against authentic experimental workflows.
●
healthThis provides a more realistic measure of how AI can assist in laboratory automation and protocol design.