Skip to content
Open access

MiniUAV-VLA: A Compact Vision–Language–Action Model for Cooperative Multi-UAV Search and Elimination via MARL Expert Distillation

Jul 2026 · Drones · 0 citations · 17 references

Abstract

Coordinating multiple unmanned aerial vehicles (UAVs) for cooperative missions requires agents that perceive their environment, reason about objectives, and generate joint actions. Vision–language–action (VLA) models unify these capabilities but lack a principled source of multi-agent training data and suffer from a training–inference discrepancy in closed-loop control. We propose MiniUAV-VLA, a compact centralized VLA controller for simulated multi-UAV search-and-elimination based on multi-agent reinforcement learning (MARL) expert distillation. A QMIX expert policy achieving 100% mission success generates multimodal demonstrations pairing rendered tactical map images with structured textual state prompts. A 158 M-parameter VLA model with approximately 65 M trainable parameters in the MiniMind-3V backbone and vision projection is fine-tuned with a multi-agent discrete action head that jointly predicts actions for all UAVs in a single forward pass. We identify a training–inference feature mismatch in behavior cloning and address it via prompt-end action pooling, which extracts action-relevant hidden states at the user–prompt boundary rather than after the generated response. In closed-loop evaluation with four drones and six mobile targets averaged over five evaluation seeds, MiniUAV-VLA reaches 74.4 ± 4.6% mission success against 9.4 ± 2.1% for a random policy and 16.2 ± 3.2% for an observation-limited greedy baseline. Across five independent training runs, prompt-end action pooling improves mean closed-loop success from 40.6% to 76.2% over the last-token alternative. These results support MARL expert distillation as a data-efficient route to compact multi-agent VLA control in this simulated setting.

Read PDF