Keyword Benchmarks Overestimate Tool-Use Capability in Small Models
October 2, 2026
Keyword-matching benchmarks produce false positives by crediting models for tool use they do not actually perform. A 1.1B parameter model scored nearly identically to a 661.6M parameter model on lenient metrics, but failed verbatim reproduction checks due to erased special tokens during web-heavy training.
HOW THIS AFFECTS YOU
●
builderDo not rely on high benchmark scores for small models when deploying tool-calling features in production.
●
researcherYou must use strict diagnostic ladders rather than simple keyword matching to validate tool-use claims.