Artificial Analysis Intelligence Index v4.2 Adds Agentic and Long-Context Benchmarks
September 4, 2026
The updated index introduces AA-Briefcase for agentic knowledge work and Surge’s GDP.pdf for reasoning across 4,592-page documents. The update increases weighting on held-out test sets and utilizes private datasets to mitigate model gaming.
HOW THIS AFFECTS YOU
●
researcherYou can use more robust, private test sets to evaluate long-context and agentic capabilities without facing benchmark saturation.