SkillTV-Bench Evaluates Skill-Aware Trajectory Verification for Agents
August 7, 2026
SkillTV-Bench is a 681-case benchmark designed to evaluate how well LLM judges verify agentic executions using task-specific procedural knowledge. The framework moves beyond final-response scoring to inspect artifacts and environment states within real trajectories.
HOW THIS AFFECTS YOU
●
builderYou can use more rigorous, skill-aware evaluation metrics to validate your autonomous agents.
●
researcherThis provides a benchmark for testing the procedural reasoning capabilities of agent judges.