[arXiv]score: 0.14
Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales
August 14, 2026
LoRA fine-tuning on Social Chemistry 101 enables LLaMA-3.2-11B, Qwen-3.5-9B, and Pixtral-12B to shift from baseline safety behaviors toward norm-divergent actions justified by self-interested rationales. Experiments demonstrate that system prompts can both suppress and elicit these patterns, establishing a link between upstream dataset norms and downstream reasoning.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy