Evaluation of Transformer and Gradient Boosting Models for Indonesian Mental Health Text Classification
Abstract
This study compares the performance between traditional feature-based classification methods and transformer architectures in mapping stress, anxiety, and depression conditions in Indonesian-language mental health discourse. The task is formulated as a multi-class classification problem, where each consultation is assigned a single dominant mental health category. By implementing an integrated experimental framework on an online consultation dataset, we tested Gradient Boosting as the baseline model against two specific transformer models, namely IndoBERT and IndoRoBERTa. Experimental findings indicate that transformer-based models consistently outperform traditional approaches, with IndoRoBERTa achieving the highest accuracy of 82%. These results affirm the capability of contextual language representation in capturing complex semantic and linguistic nuances in mental health texts. Nevertheless, this study notes ongoing challenges in differentiating categories with strong semantic overlap, particularly between stress and anxiety symptoms.