●builderStick to uniform rank allocation for LoRA when performing GRPO to avoid unexpected performance degradation.
●researcherYou should reconsider adaptive parameter allocation strategies when transitioning from supervised fine-tuning to reinforcement learning.