LoRA as Oracle uses low-rank adapters to detect malicious internalization in frozen models without needing training data or triggers. By measuring the geometric alignment of adapter updates relative to frozen weights, the method identifies backdoors even when behavioral auditing fails.
HOW THIS AFFECTS YOU
●
researcherYou can use this geometric lens to audit third-party weights for latent vulnerabilities.
●
policyThis provides a technical framework for verifying the safety of deployed neural networks.