Modeling Human Gaze Behavior Using CLIP-based Bi-Encoders
August 10, 2026
Off-the-shelf language-vision encoders can robustly replicate human gaze behavior in visual world studies. By combining a CLIP-family bi-encoder with a bimodal attribution method, the approach predicts human predictive processing without generative architectures or fine-tuning.
HOW THIS AFFECTS YOU
●
researcherYou can use existing vision-language models to study human visual attention without training new architectures.