Skip to content
Review

Mental Health Disorder Detection Beyond Social Media: A Systematic Review of Available Datasets

Jul 2026 · 2 citations · 115 references
Computer Science

TL;DR

It is found that non-social media free-text based datasets are predominantly focused on English and on detecting depression, and this first comprehensive review of non-social media, free-text datasets for mental health research is presented.

Abstract

Detecting mental health disorders in a timely manner is an important societal challenge. NLP and machine learning (ML) methods used to assist with detection rely on data collected primarily from social media. However, such datasets often have sampling biases and inherent ethical and privacy issues. One avenue to overcome these limitations is non-social media data. We present the first comprehensive review of non-social media, free-text datasets for mental health research. We use the PRISMA methodology to conduct our survey and we review datasets available in multiple languages. We find that non-social media free-text based datasets are predominantly focused on English and on detecting depression. These datasets also vary in demographics, platforms, data types, annotation techniques, and methodologies. This systematic review also reveals key gaps and highlights opportunities to develop more diverse, reliable and clinically-relevant resources.

View source

Similar papers

Conference Jul 2026

Bias-Aware Systematic Review of AI-Based Mental Health Detection Using Social Media: A PRISMA and PROBAST(+AI) Analysis

This study presents a bias-aware systematic review of artificial intelligence (AI)-based mental health detection using social media data from 2015 to 2025. Guided by PRISMA 2020 and PROBAST(+AI), records from IEEE Xplore, Scopus, and Web of Science were screened from 3,861 initial records to 308 included studies. Results show a strong shift toward transformer and hybrid models, depression-focused tasks, and Twitter/X and Reddit datasets. However, the corpus also shows major methodological gaps, including absent external validation, limited reporting of class imbalance handling, low explainable AI adoption, and underreported platform sources. The review argues that future systems require transparent data provenance, bias-aware validation, explainable decision support, and privacy-preserving crossplatform evaluation before clinical or public-health deployment.

John Rover R. Sinag, Janela Reis Babaran-Sinag · 0 citations
Conference Jul 2026

Mental Health Disorder Detection in a Low-Resource Language: A Case Study on Turkish Social Media Text

Mental health disorders such as suicidal ideation, bipolar disorder, depression, and anxiety pose public health challenges in Turkey. Approximately 20% of the population is estimated to be affected by some form of mental illness. Social media is used for sharing personal experiences and opinions in Turkey. This presents an opportunity to use textual data for detecting mental health disorders. This study aims to examine the predictive performance of various machine learning classifiers for identifying mental health conditions in Turkish language texts from social media. Given the lack of structured and unstructured datasets focused on Turkish mental health, we have mined and processed textual data from the social media platform Reddit. This led to the creation of the Turkish Mental Health Disorder (TMHD) Corpus. The TMHD dataset and the machine learning experiments conducted aim to improve the understanding of predicting mental health issues in Turkey using Artificial Intelligence.

Courage Armah, Z. Erdem, Halil Bisgin · 0 citations
Review Open access Jul 2026

From Conversations to Insights: Analysing Social Networks for Early Mental Health Detection - A Systematic Review of Causal Inference and Deep Learning

It was identified that AI continues to significantly outperform humans in terms of accuracy, efficiency, and early intervention for mental health detection.

U. Nazir, Nur Shazwani Kamarudin, Mazlina Abdul Majid · 0 citations
Review Open access Jul 2026

Predicting Mental Health Conditions from Text Using Interpretable Machine Learning

Mental health disorders such as depression, anxiety, and post-traumatic stress disorder (PTSD) affect over one billion people worldwide, yet early detection remains a major clinical challenge. In recent years, text data from social media posts, clinical notes, and patient surveys has emerged as a rich source of signals for automated mental health screening. However, most existing machine learning models operate as black boxes, limiting clinical adoption. This paper presents an interpretable machine learning framework that combines natural language processing (NLP) feature extraction with explainable AI techniques — specifically SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) — to predict mental health conditions from text while providing transparent, clinically meaningful explanations. A multi- class classification task involving depression, anxiety, PTSD, and healthy controls is performed on a dataset of 19,320 labelled text samples. The proposed XGBoost model with SHAP explanations achieves 87.3% accuracy and an AUC of 0.924, while the fine-tuned BERT model achieves 91.6% accuracy and an AUC of 0.961. Experimental results demonstrate that interpretability does not significantly compromise predictive performance, enabling trustworthy AI-assisted mental health screening.

N. Thakur, D. Patil, Pushpa Choudhary · 0 citations
Review Open access Jul 2026

AI-Based Identification of Drug Use and Overdose Signals on Social Media

The rising prevalence of substance abuse and overdose incidents underscores the need for real-time public health surveillance. Social media offers valuable signals for monitoring these events; however, noisy language, slang usage, and class imbalance present significant challenges for automated analysis. To address these issues, the authors propose ATTEND, a multi-task neural network for substance classification and detection of 18 overdose symptoms, with symptom normalization to standardized MedDRA concepts. ATTEND was trained on a large multi-source corpus combining ADE Corpus V2 and the UCI Drug Review Dataset, comprising over 100,000 samples designed to emulate realistic social media communication. Experimental results show that ATTEND achieved 93.23% accuracy and 93.41% weighted-F1 for substance classification, 94.10% micro-F1 for overdose symptom detection, and 90.42% accuracy for symptom normalization, outperforming baseline multi-task models across all tasks. The framework is scalable, privacy-preserving, and suitable for real-time monitoring of drug abuse signals.

Sudhakar Kumar, Sunil K. Singh, Satvik Pathak et al. · 0 citations