The study resulted in the development of a novel six-component framework comprising Input Processing, LLM Core, Knowledge Enhancement, Context Management, Response Generation, Response Generation, and Human Feedback that successfully addressed resource scarcity through language detection and cross-lingual query understanding.
CALAMITA is conceived as a rolling benchmark, enabling continuous integration of new tasks and models, and argues that this combination offers a blueprint for other languages and communities seeking inclusive and rigorous LLM evaluation practices.
Malvina Nissim, Danilo Croce, V. Patti et al.· Italian Journal of Computati...· 0 citations
Whether contemporary LLMs can reproduce the research outcomes of a fully documented human study: a 1991 article that identified dermatophytosis (ringworm) in historical fine art was evaluated.
: In the era of digital transformation and the rapid advancement of generative artificial intelligence, the translation of idiomatic expressions has become a crucial benchmark for evaluating the cognitive and linguistic capabilities of Large Language Models (LLMs). This paper presents a detailed analysis of research conducted on a corpus of ten English body part idioms taken from the Pioneer B2 textbook used at Singidunum University. The aim of the research was to compare translations generated by the ChatGPT model with solutions from official idiomatic dictionaries, utilising Pavol Kvetko's classification and Mona Baker’s equivalence strategies as the theoretical framework. The analysis encompasses idioms of varying degrees of transparency, ranging from completely opaque to semi-idioms. The study results indicate a 90% accuracy rate in conveying meaning, alongside an unexpectedly high 60% correspondence of keywords in both languages. The research confirms that ChatGPT successfully identifies functional equivalents in the Serbian language, often prioritising the naturalness of expressions over literal translation. This work contributes to the discussion on the role of AI tools as assistants in translation and education, emphasising that while AI shows exceptional dexterity in mapping conceptual fields, human oversight remains essential for the final validation of stylistic nuances. The findings have significant applications for international scientific research, particularly in the domain of applying information technology in foreign language teaching.
Jelena Janackovic, Jovana Bošković, Jelena Mladenović· SINTEZA· 0 citations
A multi-layer taxonomy of 14 capability domains and 91 subskills across Primitive, Constructed, and Integrative layers is introduced and supports research organization, coverage audits, evaluation interpretation, and testable hypotheses for diagnosis, training, and transfer.
Shixin Fang, Jiachen Wo, Wenjuan Qin et al.· 0 citations
A novel, layered conceptual framework is introduced that organizes research in CDR across four key dimensions: User Layer, System Layer, Data Layer, and Evaluation Layer and identifies core challenges in CDR, including the lack of standardized evaluation benchmarks and limited support for ambiguous or evolving user intent.
Lisa-Yao Gan, Johanna Walker, E. Simperl et al.· Information Systems Frontier...· 0 citations