Groundedness Drift Detects Backdoor Attacks in LLM Classifiers
August 14, 2026
Groundedness Drift is a lightweight black-box auditing metric that measures if an LLM's explanation remains anchored to the input. Testing on 7B models shows it achieves higher AUROC than existing detectors when identifying OpenBackdoor-style attacks at a 5% clean-FPR budget.
HOW THIS AFFECTS YOU
●
builderYou can use this metric to detect if your model's explanations are being used to camouflage malicious backdoors.
●
policyThis provides a concrete method for auditing the safety and reliability of automated moderation classifiers.