●builderYou can improve multilingual translation quality using reinforcement learning without needing ground-truth reference pairs.
●researcherThis demonstrates the effectiveness of gated reward models and checkpoint interpolation for reference-free post-training.