[arXiv]score: 0.18
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift
July 28, 2026
Post-training methods impact model alignment non-uniformly across 15 dimensions, including safety and factuality. Supervised fine-tuning (SFT) causes substantial behavioral and representational drift compared to reinforcement learning with verifiable rewards (RLVR), which maintains higher alignment stability while improving task performance. KL regularization serves as a key mechanism to mitigate these alignment shifts.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy