VLM Transcription Faithfulness and the FaithC4 Benchmark
July 27, 2026
Vision-Language Models often rewrite imperfect text into plausible forms rather than transcribing it accurately, a behavior missed by standard benchmarks. The new FaithC4 benchmark uses perturbations in English, Chinese, and Korean to reveal that general-purpose VLMs suffer higher Word Error Rate degradation than traditional OCR.
HOW THIS AFFECTS YOU
●
builderDo not rely on general-purpose VLMs for high-fidelity document transcription without testing against perturbations.
●
researcherThe FaithC4 benchmark provides a more rigorous evaluation for VLM transcription accuracy.