Jul 2026· European Open Science Space· pp. 95-99· 0 citations
TL;DR
The Multi- Dimensional Task-Alignment Framework (MTAF) is introduced, a novel seven-criterion evaluation instrument designed to characterize the functional specialization of competing LLMs and translate benchmark performance into domain-specific selection guidance.
Abstract
The proliferation of commercially available large language models
(LLMs) has produced a competitive ecosystem in which model selection for specific
professional tasks remains insufficiently theorized. These theses introduce the Multi-
Dimensional Task-Alignment Framework (MTAF), a novel seven-criterion evaluation
instrument designed to characterize the functional specialization of competing LLMs
and translate benchmark performance into domain-specific selection guidance.
Applying MTAF to four dominant systems – ChatGPT (OpenAI/GPT-4o), Gemini
(Google DeepMind), Grok (xAI), and Claude (Anthropic) – we identify distinct
competitive profiles: ChatGPT demonstrates leading performance in code generation,
Gemini excels in multimodal and real-time grounded tasks, Grok provides unique
access to temporally current social-media-derived data and Claude exhibits the highest
reliability in long-document processing and complex instruction following. A derived
task-model alignment matrix operationalizes these findings for practical decisionmaking
across scientific research, software engineering, academic writing, and
organizational management contexts.
This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a graphical user interface (GUI) for visualizing outcomes.
AssistEM, a framework for efficient LLM adaptation to EM via principled data selection, demonstrates that selective fine-tuning not only accelerates adaptation but also improves training efficiency (requiring fewer GPU hours), enabling open-source LLMs to rival–and in some cases outperform–closed-source models.
John Bosco Mugeni, S. Lynden, Toshiyuki Amagasa et al.· International Journal of Dat...· 0 citations
General-purpose large language models (LLMs) may struggle in specialised machine translation, but the conditions under which fine-tuning improves translation performance remain unclear. This study compares full-parameter fine-tuning (FPFT) and parameter-efficient fine-tuning (PEFT) for Chinese-English political discourse translation using a purpose-built corpus and the Qwen3-14B model. Translation performance was assessed on three 50-item test sets using BLEU, ROUGE-L F1, METEOR, and BERTScore F1, together with BLEU pass-rate likelihood-ratio G2 tests, paired t-tests, and paired Cohen’s dz for item-level score differences. The results reveal a clear contrast between unseen in-domain evaluation, maximum-overlap benchmarking, and semantically related but non-fine-tuned evaluation. On Test Set A and Test Set C, neither fine-tuning strategy produced a statistically significant BLEU pass-rate advantage over the base model, and paired tests across the continuous metrics did not show consistent fine-tuning gains. On Test Set B, which was sampled from the fine-tuning corpus, both fine-tuned models substantially outperformed the base model across all four metrics, with FPFT achieving the highest scores and PEFT providing a more computationally efficient alternative. These findings indicate that textual overlap between training and deployment data, rather than broad domain similarity alone, strongly conditions the observed benefit of fine-tuning. The study offers an empirically grounded framework for selecting fine-tuning strategies in specialised machine translation.
Cross-lingual information retrieval (CLIR) for low-resource dialects remains underexplored, despite millions of speakers worldwide. This work addresses a critical gap by introducing the first information retrieval (IR) benchmark resource for Ehugbo (the Afikpo dialect of Igbo with \textasciitilde 150,000 speakers in Nigeria), constructed from a high-quality parallel multimodal corpus: 1 hour of transcribed Ehugbo Bible audio aligned with standard English translations. This parallel design enables rigorous evaluation of retrieval models across language pairs. Our benchmark reveals a surprising and counterintuitive finding: the ''Alignment Gap'', where African-centric foundation models (Serengeti, Afro-XLMR, AfriBERTA) that excel at linguistic familiarity with Igbo achieve <5% retrieval accuracy, while global models like LaBSE achieve 85% despite less exposure to the language family. Through diagnostic analysis (t-SNE visualizations, tokenization studies, error patterns), we show that regional models lack cross-lingual alignment bridges despite deep language understanding, while global models achieve language invariance through explicit translation supervision. This finding has immediate implications for the design of multilingual systems: pre-training diversity alone is insufficient for dialectal IR. We release our Ehugbo corpus, results, and evaluation splits on GitHub to enable future work on dialect-specific fine-tuning and alignment strategies for African languages.
Ukachi Agnes Eze-Mbey, V. Olufemi, A. Bahizire et al.· Annual International ACM SIG...· 0 citations
CALAMITA is conceived as a rolling benchmark, enabling continuous integration of new tasks and models, and argues that this combination offers a blueprint for other languages and communities seeking inclusive and rigorous LLM evaluation practices.
Malvina Nissim, Danilo Croce, V. Patti et al.· Italian Journal of Computati...· 0 citations
A fundamental gap between generation fluency and reasoning ability is uncovered, a pronounced ''inverse scaling effect'' where larger models can underperform in domain-specific reasoning, and systemic bottlenecks across all SOTA models are identified, exposing fundamental limitations of current architectures.
Guangtao Nie, Huimu Wang, Gewei Lu et al.· Proceedings of the 32nd ACM...· 0 citations