Jul 2026· International Research Journal on Advanced Engineering Hub (IRJAEH)· Vol 4, pp. 4918-4926· 0 citations· 14 references
TL;DR
This work introduces the Language Model Council (LMC), a collaborative framework that combines the expertise of multiple specialized AI agents to evaluate a user query from different perspectives and outperforms traditional single-model systems by improving response quality, reducing hallucinations, and increasing user trust through enhanced explainability.
Abstract
The rapid advancement of Large Language Models (LLMs) has significantly expanded the capabilities of Artificial Intelligence in language understanding, reasoning, and automated decision support. Despite these achievements, systems built around a single language model remain vulnerable to problems such as factual inaccuracies, hallucinated information, inconsistent outputs, limited explainability, and unintended bias. These shortcomings restrict their use in applications where decisions must be accurate, transparent, and accountable. This work introduces the Language Model Council (LMC), a collaborative framework that combines the expertise of multiple specialized AI agents to evaluate a user query from different perspectives. Their independent analyses are consolidated through a consensus-driven mechanism that selects the most reliable response. To further improve transparency, the framework integrates Explainable Artificial Intelligence (XAI), providing confidence estimates together with concise reasoning summaries that clarify how the final decision was derived. Experimental evaluation indicates that the proposed approach outperforms traditional single-model systems by improving response quality, reducing hallucinations, and increasing user trust through enhanced explainability.
This study contributes to the field of artificial intelligence by offering a structured approach to building, testing, and refining multi-agent architectures that balance knowledge grounding, perspective modelling, and reasoning validation.
This research benchmarks human evaluations against a large language model (LLM) using a multi-agent approach and/or retrieval-augmented generation (RAG) to automate complex content analysis tasks to leverage artificial intelligence’s efficiency and precision alongside humans’ contextual understanding and domain expertise.
Xinyu Fu, Chaosu Li· Journal of Planning Educatio...· 1 citation
This thesis introduces the Multi-Agent LLM (MALLM) framework, which implements and evaluates various decision protocols, namely voting, consensus, and judge decision mechanisms, to simulate multi-agent discussions for conversational task solving and indicates that consensus protocols excel in knowledge-intensive domains while voting and judge protocols are more effective for logic-based tasks.
Structured multi-agent debates among Large Language Models (LLMs) have emerged as a powerful paradigm for enhancing reasoning reliability and argumentative coherence. Motivated by the European Space Agency’s (ESA) interest in trustworthy AI for space operations, this study proposes a moderated, domain-adaptive multi-agent debate framework applied to the high-stakes domain of satellite communications (SatCom). Specifically, it assesses (i) the efficacy of structured deliberation against single-agent baselines, and (ii) the impact of model heterogeneity versus homogeneity. A single-agent baseline is compared against a multi-agent framework deploying a moderator and two domain-specialized experts. These systems utilize local 70B-parameter LLMs in homogeneous (Llama-3.3) and heterogeneous (Llama-3.3 + DeepSeek-R1 + Qwen-2.5) configurations, all augmented with a shared, curated Retrieval-Augmented Generation (RAG) corpus combining academic institutional sources and ESA material from the Nebula portal (SatNex V programme). Outputs from 213 technical queries are evaluated via LLM-as-a-judge across three phases: baseline proficiency, strategic reasoning, and executive readiness. Single-agent systems lead in encyclopedic tasks, where retrieval suffices over deliberation. However, both multi-agent configurations outperform in strategic reasoning, with heterogeneous debates achieving superior performance in executive scenarios by victory margins of up to 2.75 points on a 10-point scale. These results validate architectural diversity as a decisive factor in resolving high-complexity technical conflicts. Ultimately, this work delivers a generalizable, fully traceable deliberation framework suitable for real-world, mission-critical environments. Code, prompts, and evaluation data are publicly available at
https://github.com/amozo-es/multi-agent-debate/
.
Susana Gómez Álvarez, Alejandro Mozo Quesada, Tomás Navarro et al.· Journal of Intelligence and...· 0 citations
Experiments reveal that models with similar end-to-end accuracy can exhibit markedly different agentic capability profiles, demonstrating that process-level evaluation is crucial for interpreting the true potential of LLMs and guiding the development of next-generation mathematical agents.
Jiayi Kuang, Yinghui Li, Yun-Ze Song et al.· 0 citations
An organizing framework for understanding LLM‐based agents is established, systematically deconstructing both single‐agent and multi‐agent systems into their core components, and the architectural principles and key mechanisms that underpin their intelligence are analyzed.
Yuheng Cheng, Ceyao Zhang, Zheng-Wen Zhang et al.· WIREs Data Mining and Knowle...· 1 citation