Skip to content

Author

Cornelia Caragea

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

Benchmarking large language models for target-specific financial stance detection in 10-K MD&A sections and earnings call transcripts

Financial disclosures contain rich narrative information about firm performance, but their length, specialized terminology, and target-dependent language make sentence-level analysis challenging. Existing financial sentiment methods typically assess overall positive or negative tone, rather than stance toward specific financial targets. In this study, we introduce a sentence-level benchmark for target-specific financial stance detection in Form 10-K Management’s Discussion and Analysis (MD&A) sections and quarterly earnings call transcripts (ECTs). The benchmark focuses on three financial targets–debt, earnings per share (EPS), and sales–and assigns each target-relevant sentence a positive, negative, or neutral stance label. The corpus is drawn from five public companies. The study is intended as a benchmark for NLP methods and does not support broad claims about financial disclosures in general due to the limited set of companies and financial targets. For scalability, the training split is labeled using ChatGPT-o3-pro, while the held-out test split is independently annotated by human annotators and adjudicated to form a human-consensus gold standard. Using this benchmark, we evaluate four contemporary large language models under zero-shot, few-shot, Chain-of-Thought, and document-context prompting conditions. Model outputs are assessed using accuracy, macro-averaged and weighted precision, recall, and F1, with paired statistical tests used to evaluate performance differences. Results show that LLMs can provide useful baselines for low-label, target-specific financial stance detection, but performance varies across models, document types, targets, and prompting strategies. GPT-4.1-mini and Gemma4-31B generally achieved the strongest overall performance, followed by Llama 3.3-70B and Mistral Small 3.2-24B, although no single model or prompting strategy dominated across all settings. The no-document-context condition achieved the highest Macro-F1 score. However, this result is best interpreted as alignment with the sentence-level annotation protocol, in which annotators did not have access to the full document context, rather than as evidence that document context is generally unhelpful. These findings highlight both the promise and the remaining limitations of LLM-based financial stance detection, particularly the need for broader industry coverage and economic validation.

Nikesh Gyawali, Doina Caragea, Alex Vasenkov et al. · 0 citations