Skip to content
#edge computing Preprint

Traffic-Adaptive Per-Hop Multipath Routing in Multi-Hop UAV Networks

Aug 2026 · 0 citations · 33 references
Computer Science

TL;DR

This work develops a multi-agent reinforcement learning (MARL) algorithm, termed Multi-Agent Proximal Policy Optimization with Dirichlet Modeling (MAPPO-DM), which follows the centralized-training-and-decentralized-execution framework and models continuous traffic-splitting actions using a Dirichlet distribution.

Abstract

In uncrewed aerial vehicle (UAV)-relayed mobile edge computing (MEC) networks, computation tasks generate traffic with diverse latency requirements and data sizes. Routing decisions therefore need to adapt to both traffic characteristics and changing network conditions. Compared with single-path routing, multipath routing is better suited to such heterogeneous traffic because it provides multiple forwarding options and enables flexible traffic splitting. However, conventional multipath routing usually splits traffic over predefined end-to-end paths, making it difficult to respond quickly to link fluctuations and topology changes in UAV networks. To address this issue, we propose a traffic-adaptive per-hop multipath routing method for multi-hop UAV networks, in which each UAV dynamically distributes traffic among multiple candidate next hops. We formulate the routing problem to improve the on-time packet delivery ratio while reducing the packet loss ratio, and model it as a decentralized partially observable Markov decision process (Dec-POMDP). To solve this problem, we develop a multi-agent reinforcement learning (MARL) algorithm, termed Multi-Agent Proximal Policy Optimization with Dirichlet Modeling (MAPPO-DM). MAPPO-DM follows the centralized-training-and-decentralized-execution framework and models continuous traffic-splitting actions using a Dirichlet distribution. Simulation results show that MAPPO-DM outperforms the baseline methods and maintains robust performance under various network conditions.

View source

Similar papers

Open access Aug 2026

Enhancing Routing Efficiency in UAV-Assisted Vehicular Networks Via Integrating Fog Computing and Software-Defined Networking

Unmanned aerial vehicles (UAVs) have been used in heterogeneous vehicular networks to enhance performance on extremely congested roads and areas with low coverage. Nevertheless, when aerial relays are added to the routing process, the routing occurred in a more complex environment. Routing protocols often favour UAV relay routes because UAV relay route can have better link quality and a small number of hops, but the routes developed from these routing protocols can lead to load imbalance between the aerial and terrestrial elements of the network and sometimes the UAV can be the bottleneck itself. This paper aims to utilize modern networking paradigms, i.e., Software-Defined Networking (SDN) and Fog Computing—to achieve routing operations in a heterogeneous, cluster-based Vehicular Ad Hoc Network (VANET). Fog nodes will take responsibility for offloading/performing the computational tasks involved in cluster formation, inter-segment routing between the aerial and terrestrial paths, and determining the optimal number of cluster heads. Fog nodes will use fuzzy logic and reinforcement learning to execute these tasks. The role of the SDN controller will be to manage traffic flow across fog cells using its global view of the multi-tiered network architecture which integrates heterogeneous vehicles with fog-layer connectivity. The proposed model was assessed visa a variety of routing protocols designed for UAV (Unmanned Aerial Vehicle)-assisted networks as well all routings used in traditional vehicular networks in several scenarios. The performance has proven to be far superior in a variety of aspects, including: the packet delivery ratio as a function of vehicle density and the aerial relay density; network utilization efficacy as a function of the harvesting node speed; and end-to-end delay as a function of ground node density. Finally, the results provide strong evidence on the success of the selective clustering method taken up in our model, as based on the dwell time of the cluster.

Saif Thamer Mohammed Museedi, Hardik Joshi · 0 citations
2026

Reliability and Traffic Aware Resource Allocation for UAV-Assisted Vehicular O-RAN

The rapid advancements of next-generation vehicular networks require intelligent, low-latency, and efficient resource management to support heterogeneous services. In this work, we propose a Traffic-aware Dynamic Resource Allocation (TADRA) architecture for UAV-assisted vehicular O-RAN to address the challenges of dynamic traffic conditions, infrastructure failures, and stringent quality of service (QoS) requirements. Due to the dynamic mobility and flexible deployment characteristics, UAV Open Radio Units (O-RUs) in the TADRA architecture support the terrestrial infrastructure under overload or failure conditions, dynamically extending coverage, balancing traffic loads, and restoring service to maintain uninterrupted QoS across diverse and heterogeneous traffic demands. Unlike existing static or single-layer solutions, our proposed TADRA integrates RAN Intelligent Controllers (RICs) with a Hierarchical Traffic-Aware Multi-Agent Twin-Delayed (TMT) algorithm to optimize the allocation of computation and radio resources. This joint optimization problem is NP-hard, highly dynamic, and coupled across agents, making TMT a tractable and adaptive alternative. This hierarchical framework performs traffic prioritization at the upper (application) layer and resource allocation at the lower (MAC) layer, facilitating adaptive decision-making under diverse vehicular traffic patterns. Numerical results demonstrate that our solution provides substantial gains over MATD3, MADDPG, and GA, achieving 17% lower latency, 10% higher throughput, 14% lower energy consumption, and 6.5% higher reliability.

Hayla Nahom Abishu, Ahmed Badawy, Amr Mohamed et al. · 0 citations
2026

Agentic and Embodied UAV Relays for Satellite–Aerial Networking: End-to-End Latency-Aware Optimization

Integrated satellite–aerial networks (ISANs) are emerging as a promising architecture that combines high-throughput inter-satellite transmission with the agility of uncrewed aerial vehicles (UAVs) to support flexible and low-latency traffic delivery. Owing to the inherently uneven traffic distribution in the satellite layer, traffic flows often suffer from congestion and excessive multi-hop forwarding delays. UAVs can act as adaptive relays to offload congested traffic and mitigate routing detours, thereby reducing end-to-end latency. However, latency-aware traffic management in ISANs is fundamentally challenged by highly dynamic satellite topologies, heterogeneous link characteristics, and the tight coupling between satellite traffic dynamics and UAV mobility. Existing approaches often suffer from cross-layer misalignment between satellite routing and aerial relaying, which limits coordinated latency adaptation. To address these challenges, this paper proposes an agentic UAV-assisted relay framework, termed DUS-SACUD, in which an autonomous UAV acts as an embodied agent that proactively steers traffic. First, a graph-conditioned diffusion model is developed for generative UAV–satellite link (USL) selection under dynamic network states. Second, a soft actor–critic-based reinforcement learning scheme is employed for embodied UAV deployment to minimize USL-induced delay. Through closed-loop alternating execution, DUS-SACUD jointly optimizes connectivity adaptation and mobility control in ISANs. Extensive simulations based on a realistic satellite constellation demonstrate significant end-to-end latency reduction over existing routing and UAV-assisted baselines, while maintaining robust performance under diverse ISAN conditions.

Xintong Li, Feng Wang, Qi Wu et al. · 0 citations
Conference Jul 2026

RoutePPO: eBPF-Based Proximal Policy Optimization for Adaptive Routing in UAV Swarm Networks

Unmanned Aerial Vehicle (UAV) swarm networks demand routing protocols that adapt continuously to rapid topology changes, node mobility, and fluctuating link quality. AODV may incur route-discovery overhead after topology changes, while OLSR relies on periodic topology dissemination that may lag behind fast link-quality changes; both can struggle under UAV swarm dynamics. We present RoutePPO, a closedloop adaptive routing framework that couples Proximal Policy Optimization (PPO) with eBPF-based real-time link telemetry and a P4 programmable data plane. RoutePPO-Adapt introduces a Top-K path encoder with fixed-order slot assignment and a 3-step slot-history observation, producing a topology-agnostic 30-dimensional state representation. Training uses a 9-scenario curriculum with anticipatory reward shaping and cosine learningrate decay. Across 15 deterministic routing scenarios, RoutePPOAdapt achieves a mean reward of 0.674–9.2% above the two-path baseline (RoutePPO-Base) - winning 11 of 15 scenarios while reducing latency by 31.6% and packet loss by 37.1%. A kernel-native evaluation (Linux netns + eBPF TC egress) confirms non-zero telemetry counters (0.13-0.27 Mbps), demonstrating end-to-end viability of the eBPF-PPO pipeline.

Nazım Cürmen, F. Okay, Suat Özdemir · 0 citations
Open access 2026

QEGT-Based Adaptive Routing for Energy-Efficient and Reliable Communication in UAV Swarm Networks

This study proposes an intelligent Q-learning-enhanced Evolutionary Game Theory (QEGT) routing mechanism for USNs that leverages game-theoretic incentives and Q-learning to adaptively select strategies.

Anita Murmu, Saurabh Kumar Srivastava, Nuthan Chingeetham et al. · 0 citations
Open access 2026

RACER: Real-Time Adaptive Congestion-Aware Emergency Routing in Urban Vehicular Networks

This article proposes a new approach to routing, termed RACER (Real-time Adaptive Congestion-aware Emergency Routing), which dynamically responds to changing traffic conditions without requiring additional traffic-signal-control infrastructure, relying instead on congestion information obtained through standard vehicle-to-infrastructure telemetry.

Harinath Ankarboina, Jasmini Kumari, Amit Kumar Singh et al. · 0 citations

Related blog posts

Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.