●builderWhen using VLMs for automated annotation, be aware that irrelevant visual context can destabilize judge reliability.
●researcherVLM evaluation frameworks must be carefully designed to prevent visual bias from corrupting linguistic performance metrics.