Transformers Develop Internal State Representations for Tower of Hanoi
August 10, 2026
Mechanistic interpretability shows small Transformers trained on solution traces develop a geometrically faithful Sierpinski triangle representation of the Tower of Hanoi state space. Testing on frontier models like Qwen3.6-27B reveals that while large reasoning models solve puzzles, they struggle with complex state transitions.
HOW THIS AFFECTS YOU
●
researcherThis provides evidence of emergent, decodable world models in small-scale architectures used for mechanistic analysis.