Identifying Sparse AI-Text Detection Neurons in BERT
September 28, 2026
Mechanistic analysis reveals that fewer than 1% of BERT neurons drive AI-text detection accuracy across various generators. Sparse probing and bidirectional activation patching on the RAID benchmark confirm these neurons are causally relevant to detection predictions.
HOW THIS AFFECTS YOU
●
researcherThis demonstrates that detection signals are highly localized within specific, sparse neural subsets.
●
policyUnderstanding the mechanistic basis of detection can help in evaluating the robustness of AI-generated content classifiers.