It is found that prompting models to reason in English while keeping questions in low-resource languages substantially improves UE performance, suggesting that comprehension of low-resource languages is largely intact, and that the reliability bottleneck lies in generation rather than understanding.
Andrea Alfarano, Andrea Bacciu, Saab Mansour et al.· 0 citations
This work introduces QUORUM (QUality-Optimized Routing Using Multiple annotators), a budget-aware routing framework that dynamically assigns each instance to human or LLM annotators under a fixed annotation budget and supports multiple annotations per instance, combining them through agreement-based rewards to improve reliability.
Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu et al.· 0 citations
This work formalizes Active Testing in NLP and conducts an extensive benchmarking of existing approaches across 18 datasets and 4 embedding strategies spanning 4 different NLP tasks, revealing variations in method effectiveness across different data characteristics and task types.
Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu et al.· arXiv.org· 1 citation