GradCuit optimizes instance-specific continuous states at a selected Transformer layer during test time. It uses causal self-attention to provide a differentiable path for reward-weighted gradient flow, enabling better credit assignment.
HOW THIS AFFECTS YOU
●
researcherYou can optimize latent reasoning trajectories without updating model parameters via gradient-through-circuit methods.