Skip to content

Evaluating fine-tuned, embedding-based, and zero-shot models for aspect-based sentiment analysis in South Slavic news

Sep 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 89 references
Medicine

Abstract

In the contemporary digital media landscape, the ability to automatically distill public opinion from a vast and continuous stream of information is highly important. Aspect-Based Sentiment Analysis (ABSA) offers this granular capability. In this work, we address a specific, industrially relevant formulation of this task, more formally known as document-level entity-targeted sentiment analysis (TSA) or entity-level sentiment analysis (ELSA), where the “aspects” are named entities such as companies and brands. However, its application to morphologically complex, less-resourced languages like those in the South Slavic family, particularly within the long-form news domain, remains a significant challenge. In this work, we introduce AspectBench, a new benchmark dataset for document-level ABSA, comprising real-world online news articles in Slovenian and Serbo-Croatian, designed to test generalization to unseen aspects, with the Serbo-Croatian portion being made publicly available. Using this benchmark, we conduct a comprehensive empirical study evaluating five distinct modeling paradigms, including lightweight document-embedding-based classifiers, fine-tuned multilingual and language-specific pretrained language models (PLMs), hierarchical attention networks (HANs), local large language models (LLMs) in zero and few-shot settings, and learning-to-defer (L2D) policies with multi-expert PLMs, complete LLM deferral, and confidence-gated selective LLM deferral using DSPy-calibrated prompts. Our experiments show that supervised expert models remain the most reliable foundation for this task, with Longformer, mDeBERTa-v3, mT5, language-specific encoders, and hierarchical models providing strong performance depending on language and class balance. Aspect masking is generally useful for supervised expert models, especially the PLM and HAN variants, but is not uniformly beneficial across all model families and evaluation views. Standalone local LLMs are not competitive, and complete LLM deferral is unstable when applied to every instance. However, selective LLM deferral is more promising because routing low-confidence expert predictions to the LLM can improve or stabilize performance while limiting unnecessary LLM calls. We additionally use minimum viable set analysis to show how performance, calibration, and uncertainty scale with annotation volume, providing practical guidance for future media-monitoring annotation campaigns on real-world datasets. Our work therefore fills a crucial gap for less-resourced ABSA and offers practical insights into the trade-offs between model complexity, data characteristics, and the development of robust, learning-to-defer systems for real-world media monitoring.

Read PDF

Similar papers

#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Conference Open access Dec 2013

Affordable and Energy-Efficient Cloud Computing Clusters: The Bolzano Raspberry Pi Cloud Cluster Experiment

The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.

P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al. · 110 citations · ⚡7
#computer vision Review Open access May 2017

Are Software Startups Applying Agile Practices? The State of the Practice from a Large Survey

The findings show that speed related agile practices are used to a greater extent in comparison to quality practices, and that software startups who adopt the Lean Startup approach do not sacrifice quality for speed more than other startups do.

Jevgenija Pantiuchina, Marco Mondini, Dron Khanna et al. · 84 citations · ⚡4

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.