VLM Vulnerability to Multi-Layer Typographic Decoys
September 28, 2026
Evaluation via the DecoyBench dataset shows that current vision-language models struggle to read text layered with different typographic styles. While models can extract high-contrast contour text, they almost never successfully extract text hidden behind soft shading, unlike human observers.
HOW THIS AFFECTS YOU
●
researcherThis highlights a specific robustness gap in VLM perception regarding complex visual occlusions.
●
designerBe aware that multimodal UIs relying on OCR may fail when text overlaps with complex background patterns.