GRPO Reduces Constraint Awareness and Knowledge Retention in Llama 3.1
September 22, 2026
Fine-tuning Llama 3.1 8B via GRPO increases behavioral compliance from 4% to 90% but reduces explicit constraint reporting and third-person knowledge more severely than SFT. The reward-based signal leads to context-independent token suppression rather than explicit rule adherence.
HOW THIS AFFECTS YOU
●
builderYou may encounter unexpected knowledge erosion when applying GRPO to enforce strict output constraints.
●
researcherThe findings suggest RL-based alignment may trade off internal model knowledge for surface-level compliance.