A new training method enables Multimodal Large Language Models (MLLMs) to use zoom-in tools efficiently without supervised fine-tuning. By using an InfoNCE-style reward with a curriculum of hard negative tool calls, the approach achieves competitive performance on HRBench and MME-RealWorld.
HOW THIS AFFECTS YOU
●
builderYou can implement high-resolution visual reasoning in MLLMs without expensive labeled datasets.
●
researcherThe contrastive curriculum approach provides a way to learn tool-use intrinsic rewards.