Skip to content
Preprint

Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System

Jul 2026 · 0 citations
Computer Science

TL;DR

This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a graphical user interface (GUI) for visualizing outcomes.

Abstract

The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their response quality. Prevailing methods, often confined to singular dimensions, fall short of capturing the full spectrum of model capabilities. This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a graphical user interface (GUI) for visualizing outcomes. Evaluations on the TruthfulQA dataset unveil mainstream LLMs'strengths in reasoning tasks (peaking at a composite score of 0.6104) alongside pervasive limitations in navigating complex facts and ambiguities. Transcending the narrow lens of traditional metrics, this framework offers a transparent, adaptable avenue to illuminate model potential and deficiencies. Though presently focused on English tasks, its horizons beckon toward multilingual domains. This work carves a novel path for knowledge engineering and model refinement.

View source

Similar papers

Open access 2026

Challenging the Abilities of Large Language Models in Italian: a Community Initiative

CALAMITA is conceived as a rolling benchmark, enabling continuous integration of new tasks and models, and argues that this combination offers a blueprint for other languages and communities seeking inclusive and rigorous LLM evaluation practices.

Malvina Nissim, Danilo Croce, V. Patti et al. · 0 citations
Open access Jul 2026

Multi-Criteria Evaluation of Hierarchical Reasoning, Self-Correction, and Factual Consistency in Large Language Models across Complex Language Tasks

The rapid proliferation of large language models has necessitated the development of robust evaluation frameworks that extend beyond simple accuracy metrics. This paper introduces a comprehensive multi-criteria evaluation methodology designed to assess the capabilities of these advanced computational architectures in handling complex language tasks. Specifically, the study focuses on three foundational dimensions: hierarchical reasoning, self-correction mechanisms, and factual consistency. By systematically isolating these dimensions, the research provides a nuanced understanding of how models parse intricate problem structures, dynamically revise their internal states upon detecting errors, and maintain fidelity to established external knowledge bases. The proposed framework employs novel mathematical formulations to quantify these qualitative traits, enabling a rigorous, quantitative benchmarking process. Through extensive empirical analysis across diverse datasets, the findings reveal critical trade-offs between a model's ability to engage in deep hierarchical reasoning and its capacity to remain factually grounded. Furthermore, the evaluation of self-correction capabilities highlights persistent vulnerabilities in unsupervised revision protocols. This study contributes to the broader discourse on artificial intelligence reliability and safety by offering a structured approach to diagnosing model deficiencies, ultimately guiding the design of more resilient and dependable language processing systems

Stephanie Yam · 0 citations
Open access Jul 2026

MULTI-DIMENSIONAL TASK-ALIGNMENT FRAMEWORK FOR LARGE LANGUAGE MODELS: COMPARATIVE ANALYSIS OF ChatGPT, GEMINI, GROK AND CLAUDE

The Multi- Dimensional Task-Alignment Framework (MTAF) is introduced, a novel seven-criterion evaluation instrument designed to characterize the functional specialization of competing LLMs and translate benchmark performance into domain-specific selection guidance.

Iryna Bobreshova, Olena Lebedieva · 0 citations
Preprint Jul 2026

From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models

A multi-layer taxonomy of 14 capability domains and 91 subskills across Primitive, Constructed, and Integrative layers is introduced and supports research organization, coverage audits, evaluation interpretation, and testable hypotheses for diagnosis, training, and transfer.

Shixin Fang, Jiachen Wo, Wenjuan Qin et al. · 0 citations
Preprint Aug 2026

MGAL: A Multilingual Granularity-Aware Long-Context Benchmark

MGAL is the first multilingual, granularity- and position-aware long-context benchmark, constructed from United Nations reports spanning 8K to 128K tokens across the six official UN languages, and finds that LLMs perform well at word-level tasks but struggle with coarser-grained ones.

Chunhan Li, Chenglin Xu, Zongyang Zhang et al. · 0 citations

Have Large Language Models Improved Research Methodology?

Whether contemporary LLMs can reproduce the research outcomes of a fully documented human study: a 1991 article that identified dermatophytosis (ringworm) in historical fine art was evaluated.

Fredric Narcross, Robert Marks · 0 citations