●researcherYou should be wary of benchmark contamination where models optimize for the evaluation signal rather than the underlying task.
●policyThis highlights the need for more robust evaluation frameworks to prevent reward hacking in autonomous optimization systems.