LLM Reasoning Performance Collapses Beyond Depth 10 in ProloNg Benchmark
October 9, 2026
The ProloNg testbed reveals that frontier LLMs struggle with long-horizon deductive logic using Prolog, with reasoning depth up to 22 and 62k context lengths. Most evaluated models approach chance levels once reasoning depth exceeds 10, regardless of total context window size.
HOW THIS AFFECTS YOU
●
researcherYou should account for reasoning depth, not just context length, when evaluating logical capabilities.