Low-quality instructions act as a bottleneck in preference learning by restricting response quality distributions. The proposed pipeline uses reward signals to identify weak instructions and employs rubric-guided LLM feedback to revise them, enhancing alignment without discarding data.
HOW THIS AFFECTS YOU
●
builderThis offers a method to refine existing preference datasets to improve model alignment.
●
researcherYou can improve alignment ceilings by addressing instruction quality rather than just sample volume.