ReSA uses maximum entropy inverse reinforcement learning to recover a proxy reward model from an aligned LLM's behavior. This extracted reward allows for the creation of adversarial policies that outperform existing attacks in effectiveness and model transferability.
HOW THIS AFFECTS YOU
●
researcherYou can use this to test the robustness of latent safety rewards in alignment.
●
policyThis reveals a critical vulnerability in how current LLM alignment can be bypassed via reward extraction.