Theia generates and validates 100k multimodal disaster response captions
July 31, 2026
Theia uses Qwen3.5 4B dense and 35B MoE models to generate high-fidelity textual descriptions for the vision-only Incidents1M dataset. A Qwen3.5-9B image-blind LLM-as-a-Judge pipeline is employed to validate caption semantic anchoring, facilitating more effective data-free knowledge distillation for disaster management VLMs.
HOW THIS AFFECTS YOU
●
builderYou can use this methodology to synthesize high-quality training data for domain-specific vision-language models.
●
researcherThis approach provides a scalable way to validate multimodal datasets using image-blind LLM judges.