SWARM Multilingual Dataset for Russian Propaganda Detection
September 14, 2026
The SWARM dataset contains 2,183 multilingual search engine results annotated for Russian propaganda narratives. Benchmarks show that source-based blocklists are insufficient, while the strongest LLMs achieve a 0.73 F1 score on content-level detection.
HOW THIS AFFECTS YOU
●
researcherYou can use this dataset to benchmark content-level propaganda classifiers across languages.
●
policyThis highlights the inadequacy of simple domain blocklists for managing state-sponsored influence operations.