●researcherYou should note that attribution-based filtering is currently insufficient for removing non-semantic behavioral traits from training sets.
●policyThis suggests that content-based and attribution-based filtering may not be enough to ensure model safety against subliminal traits.