APIVIS: Improving Math Reasoning via Value-Guided Gumbel Search
October 2, 2026
APIVIS introduces a training-time framework that applies finite-budget Gumbel search to chunk-level mathematical reasoning. By combining direct and searched responses with selective supervision, the method prevents GRPO from becoming ineffective when uniform group rewards are used.
HOW THIS AFFECTS YOU
●
researcherYou can improve LLM mathematical reasoning by using this value-guided search method to preserve informative learning signals during RLVR rollouts.