●builderThis helps mitigate hallucination risks in vision-language applications where models might otherwise guess object locations.
●researcherThe use of calibrated GRPO offers a method to balance refusal capability with existing performance metrics.