It is found that non-social media free-text based datasets are predominantly focused on English and on detecting depression, and this first comprehensive review of non-social media, free-text datasets for mental health research is presented.
Abstract
Detecting mental health disorders in a timely manner is an important societal challenge. NLP and machine learning (ML) methods used to assist with detection rely on data collected primarily from social media. However, such datasets often have sampling biases and inherent ethical and privacy issues. One avenue to overcome these limitations is non-social media data. We present the first comprehensive review of non-social media, free-text datasets for mental health research. We use the PRISMA methodology to conduct our survey and we review datasets available in multiple languages. We find that non-social media free-text based datasets are predominantly focused on English and on detecting depression. These datasets also vary in demographics, platforms, data types, annotation techniques, and methodologies. This systematic review also reveals key gaps and highlights opportunities to develop more diverse, reliable and clinically-relevant resources.
This study presents a bias-aware systematic review of artificial intelligence (AI)-based mental health detection using social media data from 2015 to 2025. Guided by PRISMA 2020 and PROBAST(+AI), records from IEEE Xplore, Scopus, and Web of Science were screened from 3,861 initial records to 308 included studies. Results show a strong shift toward transformer and hybrid models, depression-focused tasks, and Twitter/X and Reddit datasets. However, the corpus also shows major methodological gaps, including absent external validation, limited reporting of class imbalance handling, low explainable AI adoption, and underreported platform sources. The review argues that future systems require transparent data provenance, bias-aware validation, explainable decision support, and privacy-preserving crossplatform evaluation before clinical or public-health deployment.
John Rover R. Sinag, Janela Reis Babaran-Sinag· 2026 International Conferenc...· 0 citations
Mental health disorders such as suicidal ideation, bipolar disorder, depression, and anxiety pose public health challenges in Turkey. Approximately 20% of the population is estimated to be affected by some form of mental illness. Social media is used for sharing personal experiences and opinions in Turkey. This presents an opportunity to use textual data for detecting mental health disorders. This study aims to examine the predictive performance of various machine learning classifiers for identifying mental health conditions in Turkish language texts from social media. Given the lack of structured and unstructured datasets focused on Turkish mental health, we have mined and processed textual data from the social media platform Reddit. This led to the creation of the Turkish Mental Health Disorder (TMHD) Corpus. The TMHD dataset and the machine learning experiments conducted aim to improve the understanding of predicting mental health issues in Turkey using Artificial Intelligence.
Courage Armah, Z. Erdem, Halil Bisgin· Signal Processing and Commun...· 0 citations
It was identified that AI continues to significantly outperform humans in terms of accuracy, efficiency, and early intervention for mental health detection.
U. Nazir, Nur Shazwani Kamarudin, Mazlina Abdul Majid· Journal of Communication, La...· 0 citations
Mental health disorders such as depression, anxiety, and post-traumatic stress disorder (PTSD) affect over one billion
people worldwide, yet early detection remains a major clinical challenge. In recent years, text data from social media posts,
clinical notes, and patient surveys has emerged as a rich source of signals for automated mental health screening. However,
most existing machine learning models operate as black boxes, limiting clinical adoption. This paper presents an interpretable
machine learning framework that combines natural language processing (NLP) feature extraction with explainable AI
techniques — specifically SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic
Explanations) — to predict mental health conditions from text while providing transparent, clinically meaningful
explanations. A multi- class classification task involving depression, anxiety, PTSD, and healthy controls is performed on a
dataset of 19,320 labelled text samples. The proposed XGBoost model with SHAP explanations achieves 87.3% accuracy and
an AUC of 0.924, while the fine-tuned BERT model achieves 91.6% accuracy and an AUC of 0.961. Experimental results
demonstrate that interpretability does not significantly compromise predictive performance, enabling trustworthy AI-assisted
mental health screening.
N. Thakur, D. Patil, Pushpa Choudhary· International Journal for Re...· 0 citations
The rising prevalence of substance abuse and overdose incidents underscores the need for real-time public health surveillance. Social media offers valuable signals for monitoring these events; however, noisy language, slang usage, and class imbalance present significant challenges for automated analysis. To address these issues, the authors propose ATTEND, a multi-task neural network for substance classification and detection of 18 overdose symptoms, with symptom normalization to standardized MedDRA concepts. ATTEND was trained on a large multi-source corpus combining ADE Corpus V2 and the UCI Drug Review Dataset, comprising over 100,000 samples designed to emulate realistic social media communication. Experimental results show that ATTEND achieved 93.23% accuracy and 93.41% weighted-F1 for substance classification, 94.10% micro-F1 for overdose symptom detection, and 90.42% accuracy for symptom normalization, outperforming baseline multi-task models across all tasks. The framework is scalable, privacy-preserving, and suitable for real-time monitoring of drug abuse signals.
Sudhakar Kumar, Sunil K. Singh, Satvik Pathak et al.· International Journal of Int...· 0 citations