KoNA Benchmark Evaluates Selective Non-Compliance in Vision-Language Models
September 7, 2026
The KoNA benchmark evaluates how vision-language models handle compound queries containing both answerable and unanswerable components. It measures component-level refusal across five categories: False Premise, Visual Inaccessibility, Universal Unknown, Task Feasibility, and Safety.
HOW THIS AFFECTS YOU
●
builderThis provides a more granular way to evaluate how your VLM handles complex, mixed-intent prompts.
●
researcherYou can use this to test if models can isolate specific invalid parts of a multimodal query.