●builderYou can use QLoRA and answer-first prompting to significantly improve a multimodal model's ability to detect visual hallucinations.
●researcherThe results suggest that adaptation via fine-tuning is more effective for grounding than prompting alone.