●researcherThis provides a more rigorous way to test long-context reasoning and world-model construction compared to standard question-answer benchmarks.
●founderThis indicates that current LLM benchmarks may be misaligned with actual reasoning capabilities, potentially masking true performance gaps.