VA-Judger addresses reward hacking in joint video-audio generation by using the VAPref-10K human-preference dataset. It optimizes for semantic and temporal coherence across text, video, and audio rather than evaluating modalities in isolation.
HOW THIS AFFECTS YOU
●
builderYou can build more coherent generative media products by optimizing for cross-modal semantic alignment.
●
researcherYou can use multi-modal preference datasets to move beyond disjoint metric optimization.