[arXiv]score: 0.24
VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?
August 12, 2026
VisEditBench introduces a benchmark of 1,395 human-annotated tasks to evaluate vision-language models on iterative visualization code editing. The dataset measures performance across two distinct workflows: feedback-guided repair using buggy charts with textual instructions, and reference-guided restyling to match target chart images.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy