A framework using MLLM workflows to detect physical and structural violations in text-to-image generation. It taxonomizes failures in object structure, anatomy, and spatial relationships that standard aesthetic metrics miss.
HOW THIS AFFECTS YOU
●
researcherYou can use this to quantitatively measure physical plausibility in generative models.
●
designerThis provides a way to benchmark the real-world reliability of generated assets.