Skip to content
Open access

Linguistic markers of deception in Arabic news headlines: A cross-corpus study of stylistic and numeric features

Aug 2026 · PLoS ONE · Vol 21 · 0 citations · 40 references
Medicine

Abstract

The rapid spread of misinformation in Arabic news headlines poses a growing challenge to digital media integrity, given the scarcity of Arabic-specific automated detection tools relative to English-centric systems. Headlines are brief and context-limited, yet their lexical and stylistic patterns encode strong cues of veracity or deception, making headline-only detection practically urgent and linguistically tractable. This study investigates automatic fake-news detection using Arabic headlines, leveraging five heterogeneous corpora and their unified combination. English sets were incorporated via neural machine translation with light post-normalization that preserves stylistic cues, yielding a heterogeneous cross-domain corpus. A systematic analysis of linguistic and stylistic indicators reveals stable asymmetries between fake and real headlines that recur across domains. We evaluate approaches from classical TF-IDF baselines to Arabic-specialized transformers, and propose a late-fusion strategy coupling transformer representations with discriminative engineered features. Transformers consistently outperform classical baselines, confirming that subword representations effectively capture semantic and stylistic regularities in short Arabic texts. Late fusion yields statistically significant improvements only on datasets with prominent numeric or temporal cues; on the unified corpus, McNemar’s exact test confirms that fusion gains are non-significant, indicating that subword encoders already internalize the surface-level cues captured by the engineered features. Even where accuracy differences are marginal, interpretable features enhance explainability.

Read PDF