Sep 2026· Journal of Official Statistics· 0 citations· 12 references
TL;DR
The empirical analysis yields four main findings: a compact CNN provides the strongest model-level operational trade-off, requiring substantially less training and inference time than the LSTM or transformer while delivering similar classification reliability.
Abstract
This paper presents a deployable pipeline for constructing a news-based sentiment index (NbSI) for official-statistics use. The index is designed as a timely complement to survey-based consumer confidence measures when releases are delayed, observations are missing, or survey collection is temporarily disrupted. The pipeline is implemented using large-scale Korean economic news, manually labeled sentence-level sentiment data, pretrained word embeddings, and three standard neural classifiers: a convolutional neural network (CNN), an long short-term memory network (LSTM), and a transformer. The empirical analysis yields four main findings: (i) a compact CNN provides the strongest model-level operational trade-off, requiring substantially less training and inference time than the LSTM or transformer while delivering similar classification reliability; (ii) marginal gains from additional labeled data flatten beyond roughly 40k sentences, suggesting diminishing returns to large-scale annotation in this application; (iii) the resulting aggregate index is stable across classifier choices, supporting the use of the computationally efficient CNN as the baseline production model; and (iv) the NbSI leads Korea’s Composite Consumer Sentiment Index (CCSI) by about one month and is most useful when survey information is delayed or unavailable for sustained periods. These findings highlight the importance of transparent validation, label quality, monitoring, and maintainability when text-based indicators are adapted for official-statistics production.
Sentimental analysis (SA) of movie reviews has become an essential means of supporting audiences, filmmakers, and investors to facilitate data-driven marketing, support audience engagement, and maximize audience development returns. Nevertheless, current SA models are yet to overcome potentially daunting problems suc...
K. Raja, Bhramara Bar Biswal, R. Prasad· Natural Language Processing· 0 citations
We introduce the CDSP (context-conditional deliberation signal pipeline), converting an investment committee's meeting transcripts into structured predictive features. CDSP segments the meeting transcripts into topical chunks, assigns asset-class context labels using a large language model (LLM), maps financial keyword...
Vivek Batra, Kris Chen, Sanjiv R. Das et al.· 0 citations
In the overall comparison, FinBERT and DeBERTa-v3-base outperform the traditional baseline, whereas ModernBERT-base does not, indicating that the preferred encoder depends on the evaluation setting.
Xian-Hua Peng· Applied and Computational En...· 0 citations
The main contribution of the proposed model is therefore not absolute superiority over large transformer models, but an improved balance between accuracy, interpretability, and computational efficiency for resource-constrained ABSA applications.
Mohammad Abu Kausar, M. Nasar, Sallam O. F. Khairy et al.· Journal of Computers, Mechan...· 0 citations
It is taken as initial evidence for market time series as an input modality in financial text classification on the task of classifying sentences from Federal Reserve communication as hawkish, dovish, or neutral.
Michael Schlee, Fabian Lukassen, C. Weißer· 0 citations
In this study, a longitudinal dataset of more than 23 million news headlines from 47 U.S.-based media outlets is used to investigate the use of large language models (LLMs) for sentiment detection. Recent developments in LLMs offer potential gains in contextual understanding, adaptability, and generalization, even thou...