S2T-RLHF Addresses Unstable Reward Training via Hierarchical Credit Assignment
July 22, 2026
S2T-RLHF introduces a granularity-aware principle for preference-based RLHF to prevent training instability. It avoids the pitfalls of overly fine-grained token-level rewards that amplify noise when preference signals are only provided at the response level.
HOW THIS AFFECTS YOU
●
researcherThis approach suggests prioritizing reward stability over maximal credit assignment precision in noisy preference datasets.