Model Performance Variability Across 66 Agent-Harness Configurations
October 2, 2026
Evaluation of 66 configurations across five models and four harnesses shows that model rankings reverse depending on the agent harness used. On Terminal-Bench 4, Claude leads GPT by 7.94 points in OpenHands but trails by 30.16 points in PI, proving vendor-specific harnesses are not always optimal.
HOW THIS AFFECTS YOU
●
builderYou must benchmark your specific agent harness rather than relying on model provider benchmarks.
●
founderThis highlights that agent performance is a product of the model-harness pairing, not just the model itself.