Skip to content
Conference

Learning-Guided Task Refinement for Multi-UAV Swarm Coordination

Aug 2026 · 2026 12th International Conference on Big Data and Information Analytics (BigDIA) · pp. 1258-1265 · 0 citations · 25 references

Abstract

Multi-UAV swarm coordination requires high-level task refinement across heterogeneous game-like operation segments with different controllers, risks, and time-dependent rewards. Existing hand-written refinement rules are difficult to tune when an intermediate action has no direct reward but changes downstream losses and completion time. This paper presents a simulation-grounded learning-guided refinement framework for swarm coordination. RAE-style symbolic refinement first enumerates symbolically feasible object options, the swarm simulator evaluates their executed operation utility, and a contextual Learn-π policy distills these evaluations into deployable method and order preferences. In a controlled EMI-density sweep, Learn-π achieves zero mean held-out regret over 30 cases by selecting strike-first or suppress-first according to continuous jam coverage, whereas fixed-rule and context-free selectors cannot. On 50 randomized held-out maps it further attains the lowest zero-rollout regret among deploy-time selectors and degrades gracefully under Gaussian sensor noise on the jam feature. Full-operation and multi-scale studies show that the learned policy completes heterogeneous tasks in the 16-UAV setting and adapts visit order when airframes are scarce.

View source

Similar papers

Open access Sep 2026

Group-Level Heterogeneous Reinforcement-Learning-Guided Jellyfish Search Optimization for Cooperative Multi-UAV Path Planning Under Dynamic and Adversarial Conditions

Cooperative multi-UAV path planning must produce coupled trajectories that respect inter-UAV separation constraints through environments with moving obstacles, wind, GPS jamming, and communication loss. Population-based metaheuristics handle the resulting nonconvex, dynamic objectives well, and reinforcement-learning (...

Nader Alotaibi, Wojdan Binsaeedan · 0 citations

Temporal coordination aware reinforcement learning for multi-agent UAV navigation in dynamic environments

T-CARE is introduced, a hybrid learning-heuristic multi-UAV coordination framework that integrates zero-shot constrained action selection with priority-aware temporal reservations and achieves 100% success, 0% collision rate, and no observed persistent starvation or deadlock.

Abhudaya Shrivastava, Z. Obradovic · 0 citations
Preprint Aug 2026

Planner-Conditioned Diffusion for Coordinated Multi-Agent Exploration

A Planner-Conditioned Diffusion Policy (PCDP) is proposed, trained on demonstrations from multiple planner styles with planner identity as an explicit conditioning input, enabling a single shared model to learn a multimodal trajectory distribution and generate diverse, controllable trajectory candidates from the same o...

M. Teo, Jeric Lew, T. Duhan et al. · 0 citations
Preprint Aug 2026

AERIS: Offline Policy Improvement for Multi-UAV Integrated Sensing and Communication

This paper proposes AERIS, an offline policy improvement framework for multi-UAV ISAC that learns from fixed flight logs under centralized training and decentralized execution and designs STAR-CRDT, an offline multi-agent RL algorithm that performs support-aware local action rectification and distills only trusted impr...

Ziyuan Wang, Yi-Fan Sui, Wei Wei et al. · 0 citations
Open access Aug 2026

Efficient Exploration-Enabled Multi-Agent Reinforcement Learning for Multi-UAV Cooperative Target Search

Multi-UAV Cooperative Target Search (MCTS) is a critical task in low-altitude sensing applications, requiring agents to efficiently explore unknown environments under complex constraints. However, traditional search methods are mostly unscalable and perform poorly in dynamic multi-UAV environments. As a promising alter...

Peng Chen, Tian-Xu Li, Wei-Xing Xia et al. · 0 citations
Review Open access Sep 2026

Consensus-based path planning for UAV swarms under multiple constraints: A review

UAV swarms are essential for emergency response, logistics, reconnaissance, and environmental monitoring, yet achieving safe and scalable path planning under dynamic conditions and complex constraints remains challenging. Unlike existing surveys that categorize algorithms by theoretical foundations, this paper systemat...

Ya-Na Lu, Lian-Peng Li, Hui Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.