●builderBe aware that safety fine-tuning may introduce fragile refusal components that are difficult to steer.
●researcherInternal refusal mechanisms are highly dependent on the specific post-training optimization used.
●policySafety alignment is not a solved problem; current methods involve significant trade-offs between robustness and capability.