Models between 0.6B and 1B parameters match 70B baselines on Psych-101
August 4, 2026
Training fourteen models from 135M to 14B parameters on the 10.7M trial Psych-101 dataset shows that scale provides minimal benefit for in-distribution human behavior tasks. While larger models generalize better to out-of-distribution task structures, models under 1B parameters can match 70B baselines on held-out data.
HOW THIS AFFECTS YOU
●
researcherYou can achieve high performance on specialized cognitive tasks using significantly smaller, more efficient architectures.