An audit of 300 frozen scenes shows that text-to-3D leaderboard rankings are highly sensitive to camera settings and caption wording. Peak configuration variance between evaluators often exceeds the variance between different generative models, making current rankings unreliable.
HOW THIS AFFECTS YOU
●
researcherYou should account for high variance in prompt and render settings when evaluating 3D generation benchmarks.
●
designerBe aware that current 3D model rankings may shift significantly based on how the output is rendered or described.