Post-Training Text-to-Image Models via Composite Rewards
October 1, 2026
Implements a post-training recipe combining Bradley-Terry preference rewards with rubric-based rewards for prompt faithfulness. This dual approach captures aesthetic preferences while providing safeguards against reward hacking in open-domain generation.
HOW THIS AFFECTS YOU
●
builderYou can improve image generation faithfulness by composing diverse reward signals.
●
designerExpect better alignment between complex text prompts and visual outputs.