A supervised lexicon-learning approach is extended to 10-K filings and their Item 1A risk-factor sections, training sentiment scores against both return and volatility labels at three levels of aggregation: sector, portfolio, and individual firm.
Abstract
Financial sentiment extraction has largely relied on news text and supervised extraction against return labels alone, leaving 10-K filings -- and volatility, the target risk disclosure is arguably best suited to informing -- comparatively unexplored. We extend a supervised lexicon-learning approach to 10-K filings and their Item 1A risk-factor sections, training sentiment scores against both return and volatility labels at three levels of aggregation: sector, portfolio, and individual firm. Across 1,383 filings from 94 Nasdaq-100 technology constituents (2006--2023), we evaluate the resulting twelve sentiment metrics on classification accuracy, correlation with realised market outcomes, and qualitative lexical content. Full-filing text produces more accurate sentiment at the sector and portfolio level for both targets, but this reverses at the individual-firm level, where the narrower Item 1A section performs better -- an effect we attribute to the interaction between document volume and the amount of independent training signal available at each level of aggregation. A Loughran-McDonald dictionary baseline is consistently, strongly negatively correlated with price at every level tested, underscoring the value of a supervised approach for regulatory disclosure text. These findings, and the design choices they motivate, establish the sentiment-generation methodology underlying a subsequent, larger-scale, multi-source system.
Financial sentiment analysis has become a standard component in news-driven stock prediction, yet it reduces rich, multi-dimensional news articles to a single polarity score. We hypothesize that financial news encodes multiple orthogonal information dimensions---event type, impact scope, temporal horizon, and semantic confidence---that sentiment alone cannot capture, and that these dimensions carry independent predictive value. To test this hypothesis, we propose a structured information extraction framework that leverages LLaMA-3.1-70B to extract six semantic dimensions from financial news. Through large-scale experiments on 41,618 news--stock pairs from the FNSPID dataset, we find that (i) FinBERT sentiment features exhibit strong predictive power under nonlinear models (F1=0.576) but substantially weaker performance under linear models (F1=0.230), revealing a highly nonlinear sentiment--return relationship; (ii) LLM-extracted structured features, while individually weaker, capture information orthogonal to sentiment, as evidenced by a 53.5% systematic disagreement rate between the two approaches; and (iii) combining both signal sources yields F1=0.600, significantly outperforming either alone ($p<0.0001$), with consistent improvements across all seven event types. Ablation experiments confirm that non-sentiment structural dimensions (event type, impact subject, time horizon, confidence) independently contribute $\Delta\text{F1} = +0.019$ beyond FinBERT alone. Feature importance analysis reveals balanced contributions from all six extracted dimensions (14--21%), demonstrating that compressing news into a single sentiment score incurs substantial information loss. Our results suggest that the sentiment--semantics decoupling in financial text is systematic and exploitable, opening a new direction for multi-dimensional financial NLP.
Daohan Zhu, Sitong Ge, Ruofei Wang et al.· 0 citations
Against the backdrop of increasingly diversified and concealed forms of stock market manipulation, approaches based solely on trading data or financial indicators face growing limitations in complex and information-intensive market environments. To assess stock market manipulation risk, this study constructs a firm–month level multi-source panel dataset by retrospectively labeling violation periods at the monthly frequency based on manipulation cases sanctioned by the CSRC (China Securities Regulatory Commission). The dataset integrates corporate disclosures, investor sentiment derived from online public opinion, and market trading characteristics. A supervised learning framework that fuses textual representations and numerical features is then employed to generate manipulation risk probabilities, supporting risk ranking and tiered screening in regulatory applications. Empirical results show that the fusion model consistently outperforms single-source baselines, achieving an AUC of 0.8811 and a PR-AUC of 0.6943, along with substantial improvements in Recall@10% and Recall@20% for high-risk screening. These findings indicate that multi-source information exhibits complementary effects in manipulation risk assessment and enables effective characterization of joint anomalies along the “information disclosure–sentiment reaction–trading behavior” chain. Theoretically, this study highlights the complementary role of heterogeneous information sources, including disclosure, sentiment, and trading-related signals, in characterizing manipulation risk. In practice, it provides a feasible data-driven pathway for risk monitoring and tiered regulatory screening.
This study investigates whether employee sentiment from Glassdoor reviews is related to stock returns. Prior research relies on raw star ratings, but we show these measures are biased due to a 2012 change in Glassdoor's review process and a subsequent upward drift in scores.
We compile over two million Glassdoor reviews for Russell 3,000 firms from 2008–2019. Using multinomial inverse regression, we estimate sentiment directly from review text to address star rating biases. We then form value-weighted stock portfolios sorted on changes in sentiment.
Portfolios formed on textual sentiment changes deliver significant risk-adjusted returns, while those based on raw star ratings are inconsistent. Our evidence suggests that textual reviews contain more reliable and predictive information than numerical scores alone and that earlier findings based on star ratings may be overstated.
Analysts and investors should be cautious when using Glassdoor star ratings in investment decisions. Textual reviews offer a viable measure of employee sentiment and firm fundamentals.
This is the first study to document the change in Glassdoor collection procedures and subsequent higher average employer ratings. We demonstrate that textual analysis provides a viable method for the examination of employee sentiment. Our approach refines the link between employee satisfaction and asset pricing and offers practical insights for market participants.
M. Becker, Alexander Cardazzi, Zachary McGurk· Managerial Finance· 0 citations
This work examines the pipeline on Russell 2000 equities under three stock-selection regimes and suggests that stock-selection regime and allocator choice matter at least as much as the sentiment model, and that separating firm-specific and macro-exposure triggers is more informative than requiring both to fire simultaneously.
Alireza Kargarzadeh, Nariman Khaledian, Navid Parvini et al.· 0 citations
This paper presents an empirical comparison of lexicon-based and Large Language Model (LLM)-based sentiment analysis for extracting market-relevant signals from social media discourse in highly volatile equity markets. Using Reddit data from r/WallStreetBets and focusing on meme stocks (GME, AMC, NOK), we construct time-aligned sentiment indicators and evaluate their relationship with market returns, with particular attention to extreme positive return events in the upper tail of the return distribution. The LLM-based approach generates multidimensional sentiment representations capturing emotional polarity, bullishness, sarcasm likelihood, and topical relevance, whereas the baseline relies on the VADER lexicon-based model. We evaluate both approaches using lead/lag correlation analysis, OLS regression, ROC-AUC-based directional classification, and a quantile-based early-warning framework. The results indicate that LLM-derived indicators provide a richer multidimensional representation and exhibit stronger asset-specific statistical structure than the lexicon-based baseline. However, their relationship with market movements remains heterogeneous across assets, suggesting that increased linguistic expressiveness does not necessarily translate into stable forecasting performance in retail-driven volatility regimes.
Financial sentiment classifiers are commonly evaluated against human labels, but strong linguistic performance does not necessarily imply economically useful return predictability. This study separates these questions through two experiments. First, we construct a unified three-class benchmark from five financial text datasets and compare TF--IDF Naive Bayes, off-the-shelf FinBERT and Financial-RoBERTa encoders, zero-shot Qwen2.5-7B, and QLoRA-adapted Qwen2.5-7B, LLaMA3-8B, and Mistral-7B models. Mistral-7B achieves the best test accuracy (0.8840) and macro-F1 (0.8771), while QLoRA raises Qwen2.5's macro-F1 from 0.7274 to 0.8615. An inverse-frequency class-weighted loss does not improve Qwen2.5. Second, we evaluate economic validity on a temporally separate 2019 Benzinga sample containing 10,637 unique headlines and 13,115 headline--stock observations for a fixed S\&P~100 universe. Model probabilities are converted into continuous sentiment scores, aggregated by stock and signal date, and aligned with next-session returns over one-, two-, three-, and five-day horizons. All seven downstream models produce positive but small mean rank information coefficients at the one-day horizon; the largest is 0.0143 for FinBERT. None of the 28 model--horizon tests remains significant after Newey--West inference and false-discovery-rate correction. Portfolio results likewise fail to establish a robust advantage for the best-performing classifiers. The findings show that QLoRA is effective for financial sentiment adaptation, while also documenting a clear gap between classification accuracy and tradable cross-sectional signals.