Edge AI inference is an important workload in 5G networks. Multiple classes of edge AI workloads often share edge infrastructure, yet each may require distinct latency and rate guarantees. 5G network slicing supports differentiated requirements on the network side, but shared GPU inference also needs compute-side guarantees after requests reach the Multi-access Edge Computing (MEC) host. We extend 3GPP network slicing with compute-side enforcement so that slice guarantees remain effective after traffic reaches the MEC host. To realize this extension, we design a GPU scheduler that combines Hierarchical Token Bucket (HTB)-based traffic conditioning with Earliest Deadline First (EDF) scheduling. Our scheduler enforces per-class assured goodput, defined as the committed rate of latency-compliant completions for each class. The GPU scheduler identifies request classes via tags, which are assigned during GTPU encapsulation at the 5G user plane. This integration preserves overall latency guarantees across both the network and compute domains of a 5G slice for AI.
Yu-Hong Shen, Wen-Ju Chiang, Hsiang-Ming Hung et al.· Proceedings of the ACM SIGCO...· 0 citations
Open Radio Access Network (O-RAN) disaggregates the traditional base station into the Radio Unit (RU), Distributed Unit (DU), and Central Unit (CU) with standardized open interfaces, enabling multi-vendor interoperability and reducing deployment costs. However, realizing per-flow network slicing at the DU while simultaneously meeting the high-throughput demands of 5G remains a significant challenge. This paper presents DPDK-DU-NS, a DPDK-enabled O-RAN DU that supports per-flow network slicing and bandwidth management in compliance with 3GPP 5G QoS flow and bearer management specifications. DPDK-DU-NS separates the control plane and user plane of an OpenAirInterface (OAI)-based DU and offloads user plane functions to Intel DPDK to achieve high-speed packet processing. We further propose Adaptive Metering, a dynamic bandwidth allocation mechanism that provides Guaranteed Bit Rate (GBR) service while fairly distributing residual capacity among active QoS flows, leveraging the Two-Rate Three-Color Marker (trTCM) algorithm. Experimental results demonstrate that the system enforces 3GPP-compliant per-flow bandwidth guarantees and limits with near-perfect fairness, ensuring robust network slice isolation. These findings, coupled with the achieved 10 Gbps line-rate throughput, validate the DPDK-DU-NS user plane as a high-performance and scalable foundation for next-generation O-RAN DU implementations.