Hierarchical attention network for multilingual fake news detection
Abstract
With the rapid increase in Internet users and the increasing dependence on social media for information, the dissemination of misleading information has become a serious concern. Our objective is to address the critical issue of detecting multilingual fake news, specifically targeting fake news detection in Indian regional languages such as Hindi and Marathi. Although most of the existing studies have focused on English news, there has been limited research on the design of models for other regional languages. Due to the lack of publicly available datasets for Marathi, we created our own dataset. We implemented a two-level hierarchical attention-based model at the word and sentence levels to emphasize informative words and sentences, thereby improving prediction. A custom attention layer was designed to assign weights to words and sentences dynamically, which improves the interpretability of the model. Furthermore, different embedding methods were evaluated to improve multilingual representation performance. Experimental results demonstrate that HAN with Bi-LSTM and FastText embeddings achieved the best performance, with 98.9% accuracy and 98.8% F1-score. Although the model performs well on English, Hindi, and Marathi datasets, its ability to generalize to other Indic languages remains unexplored. The experimental results indicate that hierarchical attention mechanisms are effective for multilingual fake news detection, particularly in low-resource multilingual settings.