New Model Benchmarking on Baba Is You Logic Puzzles
July 29, 2026
The baba-is-harbor benchmark was updated to evaluate the logic reasoning capabilities of Claude Opus 5, Kimi K3, Grok 4.5, and Gemini 3.6 Flash. The test focuses on how these latest releases handle complex, rule-based environmental manipulation.
HOW THIS AFFECTS YOU
●
builderThis helps you compare the logic and reasoning capabilities of the latest flagship models for complex task planning.
●
researcherThe benchmark provides a standardized way to evaluate zero-shot reasoning in rule-governed environments.