[arXiv]score: 0.24
CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models
July 29, 2026
CogArena provides a 13-paradigm benchmark to evaluate whether LLM cognitive scores represent stable, dimensional abilities. Testing 55 open-weight models shows a single axis explains roughly 50% of variance across tasks. Targeted scaffolding offers minimal within-grouping advantages and fails to improve prediction across different model families after multiplicity correction.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy