A factor analysis of 156 models identifies a unified general capability factor, but the new NeuroCognition benchmark reveals specific failures in abstract reasoning and spatial working memory. The benchmark uses adapted neuropsychological tests like Raven's Progressive Matrices to distinguish between task completion and foundational cognition.
HOW THIS AFFECTS YOU
●
researcherYou should look beyond task-based benchmarks to evaluate the actual cognitive architectures of models.