SMALL LANGUAGE MODELS FOR MULTILINGUAL DISINFORMATION RETRIEVAL
Abstract
Disinformation campaigns reuse known narratives while changing the language, actors, and publication context. Such variation makes it difficult for analysts to connect new material with cases already documented by fact-checkers. To address this task, a software application was developed for multilingual narrative normalization and the retrieval of documented debunks from the EUvsDisinfo corpus. The retrieval corpus contains 7,538 fact-checking records published between 2015 and 2026. Each record links a disinformation claim to its debunk or a contextual explanation. The experiment covered materials concerning Russia’s war against Ukraine, Russia–NATO relations, Syria, and COVID-19. Across these data, the same semantic patterns recur in different languages and event contexts. The language model converts an input text into a concise narrative query. It retains the main claim and the source’s position while removing names, dates, and other contextual details. The system uses the resulting query for semantic retrieval and corpus record reranking. This module can operate as part of a disinformation-monitoring system: it connects new material with known narratives and passes the retrieved debunks to an evidence-verification component. The study compares the capabilities of the commercial large language model Gemini Flash 3.5 with the locally deployed Qwen3-4B and MamayLM-Gemma-3-4B models. The results confirm that small language models can perform narrative normalization in a multilingual retrieval system with a moderate reduction in quality compared with the commercial model. Local deployment provides control over data and enables fine-tuning to improve model quality and stability.