ProgramDistill Benchmark for Verifiable Web Development Tasks
September 15, 2026
ProgramDistill evaluates coding agents by requiring them to infer functionality from existing web applications. The pipeline discovered 1,975 replay-verified behaviors and 4,063 tasks, with GPT-6 Astra and Claude Opus 5 achieving 49.2% success.
HOW THIS AFFECTS YOU
●
builderYou can benchmark agents on more realistic web development tasks rather than just isolated code snippets.
●
researcherThis offers a method for generating high-fidelity, verifiable coding benchmarks without human intervention.