VTC-Bench Evaluates LLM Output Diversity and Task Coverage
August 26, 2026
VTC-Bench introduces Validated Task Coverage (VTC) to measure how many distinct, useful results an LLM produces within k attempts. The five-domain benchmark uses real-data tasks to evaluate output diversity and distinctness without relying on model-based judges.
HOW THIS AFFECTS YOU
●
builderThis allows you to optimize your inference settings for applications where multiple candidate outputs are required.
●
researcherYou can use VTC to evaluate whether model improvements actually increase the breadth of useful candidate outputs.