Semi-Supervised Text Classification for Public Expenditure Analysis
Abstract
The verification of personnel expenditure in Brazilian municipalities is critical for ensuring compliance with the Fiscal Responsibility Law. However, inconsistencies in the classification of public expenditures - often due to errors or deliberate misreporting - pose challenges for accurate auditing. While prior studies have addressed this issue using supervised learning methods, they rely on large volumes of expert-labeled data, which are costly and time-consuming to obtain. This paper proposes a semi-supervised learning approach for classifying public expenditure texts, enabling effective model training with a minimal number of labeled instances. Using a dataset of over 17,000 expenditure records labeled by experts, we simulate a lowresource scenario with only 235 labeled examples and leverage a self-training strategy to iteratively label unlabeled data. Our results demonstrate that the proposed method achieves high classification performance (macro F1-score of 0.89 in crossvalidation and 0.86 on external data), comparable to fully supervised methods, while significantly reducing annotation costs. The methodology offers a scalable, cost-effective solution to enhance the accuracy of public accounts auditing, a domain where annotated data are scarce.