LongWoF-Bench is a new benchmark comprising 778 machine-verifiable tasks in code, math, and agent-environment synthesis. It evaluates whether models can maintain interdependent constraints and utilize externalized execution experiences through EvoMap Genes.
HOW THIS AFFECTS YOU
●
builderYou can better measure the reliability of your agents in complex, multi-step workflows.
●
researcherThis provides a framework for studying how models reuse successful execution trajectories.