●builderYou can implement rubric-based reward signals to prevent your agents from becoming overly brief or unhelpful when instructed to be accurate.
●researcherThis offers a more nuanced method for training models to be both truthful and informative during long-form generation.