DASO Method for Tree-Aware Semantic-ID Optimization in Recommendations
August 24, 2026
DASO addresses reward degeneracy in generative recommendation when using GRPO for tree-structured Semantic-ID tasks. By treating training as an online rollout-allocation problem, it prevents weak reward signals that occur when top beam candidates miss the target SID branch.
HOW THIS AFFECTS YOU
●
researcherYou can improve the efficiency of post-training hierarchical item identification by solving target-missing issues in GRPO.