●builderSubmit rate is a misleading proxy for agent reliability — you need test-verified resolve rate as a production metric, and should design for silent semantic failure detection.
●researcherSilent semantic failure — repeated identical misinterpretations rather than random errors — is a distinct failure mode that current benchmarks undercount and needs targeted evaluation.
●founderGemini outresolves GPT-5 despite submitting 30% less often, which matters for product decisions around which model to deploy in agentic coding workflows.