●builderYou might achieve better agentic performance by using base models with optimized test-time compute rather than relying solely on post-training.
●researcherThis highlights a critical trade-off between single-shot accuracy and the diversity of the solution space in RLHF.