Human testers scored approximately 48% on ARC-AGI-3 tasks according to OpenAI feedback. The GDPval benchmark is reaching saturation due to highly specified task structures that differ from real-world application requirements.
HOW THIS AFFECTS YOU
●
researcherYou should look for benchmarks beyond GDPval as task saturation may limit its utility for evaluating general reasoning.