●builderThis benchmark provides a more accurate signal for how your agents will perform in production client environments.
●researcherYou can use this to evaluate agents on their ability to generate functional software rather than just passing static tests.