OrderProbe Benchmark Reveals Low Structural Reconstruction in LLMs
August 26, 2026
The OrderProbe benchmark uses fixed four-character expressions in CJK languages to evaluate an LLM's ability to reconstruct scrambled inputs. Experiments across twelve models show zero-shot structural recovery frequently falls below 35%, indicating significant weaknesses in deterministic structural reasoning.
HOW THIS AFFECTS YOU
●
researcherThis identifies a specific reasoning gap in frontier models regarding structural versus semantic understanding.