Skip to content
Preprint

Model-Based Reinforcement Learning for Heterogeneous Multi-Robot Task Assignment Under Distribution Shifts

Aug 2026 · 0 citations · 111 references
Computer Science

TL;DR

A prediction-aware adaptive rollout framework for heterogeneous multi-robot task assignment with scheduled and real-time requests that achieves near-complete service and reduces serviced-request wait times relative to reactive, token-passing, prediction-positioning, and myopic greedy baselines.

Abstract

Heterogeneous multi-robot service systems must assign requests to compatible robots, construct feasible schedules, and adapt as new tasks arrive online. Historical data can help anticipate future demand, but relying too heavily on inaccurate predictions can degrade performance under distribution shifts. We develop a prediction-aware adaptive rollout framework for heterogeneous multi-robot task assignment with scheduled and real-time requests. The problem is formulated as a finite-horizon stochastic dynamic program incorporating robot-task compatibility, ordered service requirements, routing constraints, service windows, and end-of-horizon return requirements. The proposed policy evaluates current assignments using sampled future request scenarios while restricting immediate commitments to requests already observed. To enable online use, the framework combines pruned candidate controls, wait actions, and an interaction-aware base policy for efficient future-cost estimation. Robustness to forecast error is provided by adaptively reweighting predicted requests based on recent prediction mismatch and selectively re-optimizing assigned but unstarted requests. We also introduce a historical-data-driven procedure for selecting the heterogeneous fleet composition before deployment. In a case study using real nursing-task requests from hospital inpatient floors, the proposed approach achieves near-complete service and reduces serviced-request wait times relative to reactive, token-passing, prediction-positioning, and myopic greedy baselines, with the largest improvements in tail-delay metrics.

View source

Similar papers

Preprint Aug 2026

MARA: Flow-Matching-Guided Multi-Agent Resource Allocation for Computational Resource Efficient Learning

This work proposes MARA, which predicts future loss trajectories with conditional flow matching and coordinates compute nodes through a cooperative multi-agent autoregressive policy and reduces remaining-resource prediction error relative to weighted least squares.

Han-Ye Zhao, Mu-Ning Wen, Yong Yu et al. · 0 citations
Conference Nov 2025

Multi-Objective Task Allocation and Path Planning in Heterogeneous Multi-Robot Systems Using Hierarchical Framework and Reinforcement Learning

This study proposes a multi-objective optimization-based framework for task allocation and path planning to address the challenges faced by multi-robot systems in transport-oriented task environments. The framework considers robot capability heterogeneity and load capacity, aiming to minimize task execution time and ov...

S.-H. Lo, R. Chen · 0 citations
Open access Aug 2026

ICBBA-ACO-Based Multi-Robot Task Allocation for Smart Charging Stations

Smart charging stations require mobile charging robots to respond to dynamically arriving charging requests with heterogeneous priorities, varying travel costs, and uneven workloads while maintaining online scheduling feasibility. Conventional single-layer approaches often optimize task assignment or route ordering sep...

Meiyu Chang, Zhaoyu Ku, Xuan-Yu Xing et al. · 0 citations
Preprint Aug 2026

Shared Prefixes, Better Credit: Adaptive Routing for Multi-Agent Reasoning

Multi-agent reasoning (MAR) improves reasoning reliability through iterative solution exchange and refinement. Existing adaptive MAR methods typically learn routing decisions from query-level labels or trajectory-level returns, but such coarse supervision cannot accurately estimate the state-conditioned utility of indi...

Yi-Qing Liu, Zi-Hao Wang, Han-Tao Yao et al. · 2 citations
#artificial intelligence Preprint Sep 2026

AssemblyGrid v1: A Benchmark for Multi-Robot Production with Temporary Coalitions, Local Information, and Geometric Constraints

AssemblyGrid v1 is introduced, a reproducible benchmark for repeated multi-robot production that combines explicit process progression, decentralized observations, material transfer, temporary multi-robot coalitions, productive concurrency, and geometry-dependent feasibility within one task-level formulation.

Fouad Bahrpeyma, David Heik, Dirk Reichelt · 0 citations
Preprint Aug 2026

Multi-Agent Reinforcement Learning for Online Traffic Scheduling in Time-Sensitive Application

This paper proposes a multi-agent reinforcement learning (MARL) framework for TSN scheduling, where each TSN queue is modeled as an autonomous agent and the Heterogeneous-Agent Proximal Policy Optimization (HAPPO) algorithm is employed to explicitly model inter-agent dependencies and jointly optimize service delivery a...

Marcos Carvalho, Fatih Temiz, Shavbo Salehi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.