Reasoning Models Fail Systematicity Tests in Rule Induction
September 11, 2026
Evaluation of reasoning models on structurally equivalent rule induction tasks shows significant performance gaps between original tasks and isomorphic variants. Models frequently solve a primary task but fail on recombined or substituted versions, indicating a lack of true compositional systematicity.
HOW THIS AFFECTS YOU
●
researcherYou should account for structural isomorphisms when evaluating the robustness of reasoning capabilities.