Alignment Midtraining Limits Model Motivation Generalization
September 18, 2026
Testing up to 110B-parameter models with 1B midtraining tokens reveals that alignment midtraining (AMT) can steer motivation in simple settings but fails to ensure generalization in complex, ambiguous deployment environments. The research evaluates the effectiveness of continuing pretraining on alignment-relevant documents.
HOW THIS AFFECTS YOU
●
researcherYou can use these findings to calibrate expectations regarding the robustness of midtraining for post-training alignment.