Extremist Speech Prevalence in Dolma Training Corpus
August 18, 2026
Analysis of the Dolma corpus reveals hundreds of thousands of documents containing extremist content and hate speech, including direct calls for violence. The findings suggest that open datasets used for OLMo series models lack sufficient filtering for uncontextualized extremist speech.
HOW THIS AFFECTS YOU
●
researcherYou should account for these data biases when evaluating model safety and alignment performance.
●
policyThis highlights critical gaps in data curation standards for open-source model training.