KonkaniFood 2.0: An Explainable Code-Mixed Marathi English Multilingual Transformer-Based Dataset for Sentiment Classification
Abstract
The widespread adoption of social media platforms has increased the availability of code-mixed textual data, particularly for low-resource languages such as Marathi–English. But sentiment analysis has considerable obstacles stemming from multilingual diversity, transliteration complexity, and the scarcity of high-quality annotated datasets. This work presents `KonkaniFood 2.0`, a 5,195 YouTube comments dataset of Konkani cuisine, collected through web scraping and carefully annotated for positive, negative and neutral sentiment (3-Class), with a Fleiss's kappa score of 0.967. Sentiment categorisation was carried out using Advanced Transformer-Based Language Models: mBERT, MuRIL,and IndicBERT. MuRIL model showed a better understanding of the semantic nuances of code-mixed text with 98% accuracy. Moreover, explainable AI (XAI) approaches were employed to interpret the model predictions by highlighting sentiment-carrying phrases, hence boosting transparency and trustworthiness. This work offers a valuable resource in Marathi-English (Mr-En) resource-scarce, code-mixed multilingual sentiment analysis.