2020· International Journal of Intelligent Automation & Robotics Engineering· Vol 3, pp. 01-16· 0 citations
TL;DR
It is concluded that MARL is a promising solution for future intelligent collaborative robotics and highlights future research directions including federated reinforcement learning, explainable AI, edge-based robotic intelligence, and adaptive swarm robotics for Industry 4.0 applications.
Abstract
Multi-Agent Reinforcement Learning (MARL) has emerged as an important approach for coordinating collaborative robots in industrial automation, warehouse logistics, healthcare, autonomous vehicles, and distributed robotic systems. Traditional centralized robot coordination methods faced limitations such as poor scalability, low adaptability, synchronization issues, and weak fault tolerance in dynamic environments. MARL overcomes these challenges through decentralized learning, where multiple robotic agents interact with the environment, learn from rewards, and improve coordination strategies autonomously. Before 2019, MARL gained significant attention in applications like cooperative navigation, formation control, multi-robot exploration, task allocation, path planning, collision avoidance, and resource sharing. This survey reviews key MARL techniques including Q-learning, Deep Q-Networks (DQN), policy-gradient methods, actor-critic models, and cooperative game-theoretic approaches for robotic coordination. The study explains a structured MARL coordination framework involving environment modeling, state representation, reward optimization, agent communication, and distributed decision-making. Experimental results show that MARL-based robotic systems improve task efficiency, coordination accuracy, energy optimization, adaptability, and collision reduction compared to centralized or heuristic methods. However, challenges such as communication delays, scalability, reward sparsity, non-stationary environments, and convergence instability still remain. The paper concludes that MARL is a promising solution for future intelligent collaborative robotics and highlights future research directions including federated reinforcement learning, explainable AI, edge-based robotic intelligence, and adaptive swarm robotics for Industry 4.0 applications.
Cooperative navigation of multi-agent UAVs in complex environments faces key challenges including local optima traps, sparse rewards, learning imbalance among agents, and insufficient cross-scenario generalisation. This paper proposes a multi-agent deep reinforcement learning framework that addresses these issues through coordinated exploration, demonstration exploitation, safe curriculum scheduling, and structure-aware generalisation. First, a perception mechanism combining memory of visited states, directional novelty estimates, and penalty backpropagation enables agents to proactively detect and escape local optima. Second, a hierarchical collaborative demonstration buffer with tiered behaviour cloning manages trajectories by degree of team collaboration and applies differential supervision to the actor network, improving demonstration utilisation under sparse collaborative signals. Third, a safety-aware dual-condition curriculum scheduling mechanism reviews mastered scenarios through back-testing and experience pre-filling during training, suppressing catastrophic forgetting while ensuring both task performance and flight safety. For generalisation, local geometric features computed from sensor readings are abstracted into a domain parameter, through which a structure-aware gating network and mixture-of-experts mechanism condition the policy on local structural patterns rather than scenario-specific coordinates, enabling cross-scenario transfer without exposure to the target environment. The framework is further validated under mixed static-dynamic obstacle settings, showing robust adaptability to dynamic disturbances. Simulation results confirm strong performance in collaboration success rate, navigation robustness, zero-shot cross-scenario generalisation, and dynamic environment adaptability.
Modern manufacturing faces increasing demands for flexibility, customization, and productivity under dynamic conditions. Multi-robot systems offer a promising solution by enabling cooperative execution of complex tasks, such as assembly and cooperative manipulation. In this context, Multi-Agent Reinforcement Learning (MARL) has emerged as a promising paradigm to enhance coordination and adaptability in industrial settings. MARL enables multiple agents to learn and interact in shared environments to achieve common goals within complex and dynamic industrial processes. In this paper, a deep analysis of MARL applied to industrial multi-robot systems based on a systematic review is presented, with particular focus on cooperative manipulation tasks. Following PRISMA guidelines, we analyze a total of 30 articles published between 2016 and 2026, selected independently by two of the authors from an initial pool of 102 records retrieved from Scopus and Web of Science. These articles were used to address five key questions regarding MARL algorithms, control architectures, industrial applications and validation practices. These research questions seek to examine gaps and trends at the research level which are important for the development of multi-agent control technologies. This review shows a clear prevalence of model-free algorithms under Centralized Training with Decentralized Execution (CTDE) architectures, with validation mainly performed in simulation. Despite promising results and high potential for impact, critical gaps remain in scalability, reproducibility, and sim-to-real transfer, limiting real deployment in manufacturing environments. To address these challenges and fill current gaps, we outline actionable research directions, such as hybrid MARL approaches, standardized industrial benchmarks, digital twin pipelines, and safety-aware deployment strategies, to accelerate MARL adoption in industrial environments.
Francisco J. Huertos, Oihane Bañales, Pedro Álvarez et al.· Robotics· 0 citations
Mobile manipulators on construction sites offer considerable potential for increasing productivity, as the transport of materials and the execution of precise assembly work can be increasingly automated. However, the coordination of several such robots is a complex planning task, as task assignment, navigation, and reachability planning must be solved simultaneously and under dynamic environmental conditions. This work presents a multi-agent reinforcement learning (RL) approach that enables multiple mobile manipulators to complete a set of tasks in a structured environment. Each agent makes decentralized decisions about task selection and navigation, with kinematic reachability ensured by an integrated inverse kinematic solver. The policy is trained using proximal policy optimization (PPO), supported by a reward function that encourages both navigation progress and efficient task distribution. Simulation results show that the trained model is able to efficiently distribute tasks among multiple robots while taking kinematic constraints into account. The proposed method is superior to a greedy baseline that selects the nearest available task. With four robots and 35 tasks, the multi-agent RL approach achieves a success rate of 100%, while the baseline reaches only 55%.
Charlotte Stein, Yuheng Zhi, Michael C. Yip et al.· 2026 IEEE/ASME International...· 0 citations
This paper presents a narrative survey of recent developments in MARL and examines research directions centred on centralised training with decentralised execution (CTDE), value decomposition, learned communication, graph-based methods, and model-based learning.
A. Rakib, K. Phung, Marco Pérez Hernández et al.· Applied Sciences· 0 citations
A multi-agent guided soft actor–critic (MAGSAC) deep reinforcement learning algorithm to enable multiple UAVs to simultaneously arrive at multiple constant-velocity moving targets and outperforms existing mainstream algorithms in synchronization success rate, temporal synchronization accuracy, and safety.
Shuanli Jia, Naiming Qi, Zheng Li et al.· Drones· 0 citations