Mitigating Spurious Correlations in Mental Health Analysis: A Contrastive Learning Approach for Cross-Domain Robustness
Abstract
To address the susceptibility of pre-trained language models (PLMs) to spurious correlations in mental health analysis, we propose a robust framework designed to disentangle genuine clinical signals from superficial artifacts. Our approach integrates a contrastive learning method by using masked language modeling (MLM) with artifact-masked data and an adaptive hard-negative sampling strategy to ensure cross-domain generalization. Extensive evaluations on diverse benchmarks, including real-world social media posts and LLM-generated synthetic conversational utterance datasets, demonstrate that our method significantly outperforms baselines in out-of-distribution (OOD) settings, particularly as contextual information increases. Importantly, the learned representations capture clinically meaningful patterns consistent with well-documented psychiatric comorbidities (e.g., depression-eating disorder), suggesting genuine clinical signal capture rather than superficial pattern matching. These findings demonstrate the effectiveness of our framework in detecting mental disorders in real-world settings, suggesting its practical applicability in social media platforms and conversational AI systems.