Overall, how agents are differentiated produces larger performance shifts than topology, which should be evaluated jointly with specialization: how agents are differentiated produces larger performance shifts than topology, which should be evaluated jointly with specialization.
Abstract
Multi-agent LLM systems combine multiple inference calls, but prior work often confounds how calls are connected with how they are diversified. We study these factors independently: inference topology and source of inter-agent diversity. In a controlled $2 \times 3$ matrix, we cross parallel aggregation and sequential refinement with stochastic sampling, role prompting, and learned QLoRA specialization, under a fixed three-call budget and output protocol within each backbone. Using Qwen2.5-14B-Instruct and Llama-3.1-8B-Instruct, we evaluate all six configurations on multilingual low-resource emotion detection across nine languages. Parallel learned specialization is strongest on Qwen at 52.83 Macro-F1 and reaches 52.94 on Llama. On Qwen it also exceeds same-backbone zero-shot, few-shot, CoT, and seven-call self-consistency baselines. The preferred topology depends on diversity source: sequential refinement helps stochastic and prompted settings, while the learned Width advantage shrinks from 2.83 points on Qwen to 0.17 on Llama. Depth-wise analysis suggests that later learned specialists can overwrite correct early predictions, although the aggregate effect is backbone-dependent. Overall, how agents are differentiated produces larger performance shifts than topology, which should be evaluated jointly with specialization.
Modern IT helpdesks increasingly need to support multiple languages while staying transparent and resistant to hallucination, yet most deployed chatbots still rely on single-model architectures that struggle on both counts. This paper presents NexaServe v2, a four-tier cascading multi-agent IT helpdesk system combining...
Syed Muhammad Umar Afnan, A. Nellyet· International Journal of App...· 0 citations
The proposed Mixture of Roles (MoRe), which adaptively composes multiple specializations into a single steering vector for single-turn inference, enables multi-perspective specialization in a single-agent, single-turn inference process.
Zhichen Zeng, Hui-Yuan Chen, Jingru Cheng et al.· 2 citations
Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit assignment, where su...
Thanh-Dat Truong, Sankalp Pandey, Hugh Churchill et al.· 0 citations
With the rapid advancement of large language models (LLMs), multi-agent systems have emerged as a promising alternative to scaling up a single model. Existing approaches ensemble multiple LLMs to improve response quality, but they often rely on static prior knowledge of model capabilities and prompts, and require exten...
Jinkun Xu, Minghan Wang, Zhiyong Wang et al.· Proceedings of the 32nd ACM...· 0 citations
Automated Machine Learning (AutoML) reduces the technical requirements for applying machine learning, yet existing systems still demand technical expertise, optimize inefficiently, and rarely support conversational use. We present a conversational AutoML system with a five-agent architecture coordinated via natural lan...