Skip to content

SWARM: A Multilingual Human-Annotated Dataset for Russian Propaganda Detection in Search Engine Results

Sep 2026 · 0 citations · 71 references
Computer Science

TL;DR

SWARM (Search-Web documents Annotated for Russian propaganda, Multilingual), a dataset of 2,183 search engine results across nine languages and diverse web domains, each annotated by trained coders for whether it supports a recurring Russian propaganda narrative is introduced.

Abstract

Russian state propaganda spreads across many languages and online spaces. Yet, most computational work examines only one such space, usually social media, in one or two languages, and analyses sources rather than content. We introduce SWARM (Search-Web documents Annotated for Russian propaganda, Multilingual), a dataset of 2,183 search engine results across nine languages and diverse web domains (e.g., news, blogs, government sites), each annotated by trained coders for whether it supports a recurring Russian propaganda narrative. We benchmark a source-based blocklist, supervised classifiers, and zero-shot LLMs against these labels. The blocklist misses most propaganda-supporting documents, because such content is not confined to flagged"propaganda"outlets but also appears on mainstream ones. Content-level analysis helps, though how much depends on the model: the strongest LLM reaches a positive-class F1 of 0.73, whereas the supervised classifiers reach only about 0.5, with the smaller LLMs over-predicting support, mistaking topical relevance for endorsement. Detecting search-borne propaganda thus requires per-language, content-level evaluation, which we hope SWARM and our evaluation code enable.

View source

Similar papers

Open access 2026

LoveHate: Stance Detection and Generation for Multiple Topics in User-generated Comments in Russian and English

This paper introduces LoveHate, a new multi-topic corpus of user-generated arguments in Russian, collected from the historical data of the debate platform lovehate.ru. The dataset contains nearly 19,000 posts spanning 16 socially and politically relevant topics, each mapped to binary pro and con stances. We test multip...

Natalia Evgrafova, Veronique Hoste, Els Lefever · 0 citations
Open access Aug 2026

Multilingual Fake News Detection Using Machine Learning with Contextual-Based Feature Extraction

The proposed approach provides a simple and efficient solution for multilingual fake news detection in data-scarce environments with ensemble-based classifiers such as Random Forest and Gradient Boosting achieving reliable performance across both languages.

Nikita Garg, Pritam Singh Negi · 0 citations
Preprint Aug 2026

ProBel: Propaganda Detection with Techniques, Spans, and Explanations

Propaganda detection includes several related prediction levels, ranging from sentence-level decisions to technique classification and span identification. However, it remains unclear how supervision at these levels interacts when learned jointly across Arabic and English. We present ProBel, an Arabic and English resou...

Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Elisa Sartori et al. · 0 citations
Open access Sep 2026

SMALL LANGUAGE MODELS FOR MULTILINGUAL DISINFORMATION RETRIEVAL

Disinformation campaigns reuse known narratives while changing the language, actors, and publication context. Such variation makes it difficult for analysts to connect new material with cases already documented by fact-checkers. To address this task, a software application was developed for multilingual narrative norma...

Oleh Melnychuk · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.