This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a graphical user interface (GUI) for visualizing outcomes.
Abstract
The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their response quality. Prevailing methods, often confined to singular dimensions, fall short of capturing the full spectrum of model capabilities. This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a graphical user interface (GUI) for visualizing outcomes. Evaluations on the TruthfulQA dataset unveil mainstream LLMs'strengths in reasoning tasks (peaking at a composite score of 0.6104) alongside pervasive limitations in navigating complex facts and ambiguities. Transcending the narrow lens of traditional metrics, this framework offers a transparent, adaptable avenue to illuminate model potential and deficiencies. Though presently focused on English tasks, its horizons beckon toward multilingual domains. This work carves a novel path for knowledge engineering and model refinement.
CALAMITA is conceived as a rolling benchmark, enabling continuous integration of new tasks and models, and argues that this combination offers a blueprint for other languages and communities seeking inclusive and rigorous LLM evaluation practices.
Malvina Nissim, Danilo Croce, V. Patti et al.· Italian Journal of Computati...· 0 citations
The rapid proliferation of large language models has necessitated the development of robust evaluation frameworks that extend beyond simple accuracy metrics. This paper introduces a comprehensive multi-criteria evaluation methodology designed to assess the capabilities of these advanced computational architectures in handling complex language tasks. Specifically, the study focuses on three foundational dimensions: hierarchical reasoning, self-correction mechanisms, and factual consistency. By systematically isolating these dimensions, the research provides a nuanced understanding of how models parse intricate problem structures, dynamically revise their internal states upon detecting errors, and maintain fidelity to established external knowledge bases. The proposed framework employs novel mathematical formulations to quantify these qualitative traits, enabling a rigorous, quantitative benchmarking process. Through extensive empirical analysis across diverse datasets, the findings reveal critical trade-offs between a model's ability to engage in deep hierarchical reasoning and its capacity to remain factually grounded. Furthermore, the evaluation of self-correction capabilities highlights persistent vulnerabilities in unsupervised revision protocols. This study contributes to the broader discourse on artificial intelligence reliability and safety by offering a structured approach to diagnosing model deficiencies, ultimately guiding the design of more resilient and dependable language processing systems
Stephanie Yam· Journal of innovative resear...· 0 citations
The Multi- Dimensional Task-Alignment Framework (MTAF) is introduced, a novel seven-criterion evaluation instrument designed to characterize the functional specialization of competing LLMs and translate benchmark performance into domain-specific selection guidance.
Iryna Bobreshova, Olena Lebedieva· European Open Science Space· 0 citations
A multi-layer taxonomy of 14 capability domains and 91 subskills across Primitive, Constructed, and Integrative layers is introduced and supports research organization, coverage audits, evaluation interpretation, and testable hypotheses for diagnosis, training, and transfer.
Shixin Fang, Jiachen Wo, Wenjuan Qin et al.· 0 citations
MGAL is the first multilingual, granularity- and position-aware long-context benchmark, constructed from United Nations reports spanning 8K to 128K tokens across the six official UN languages, and finds that LLMs perform well at word-level tasks but struggle with coarser-grained ones.
Chunhan Li, Chenglin Xu, Zongyang Zhang et al.· 0 citations
Whether contemporary LLMs can reproduce the research outcomes of a fully documented human study: a 1991 article that identified dermatophytosis (ringworm) in historical fine art was evaluated.