Skip to content
Book Open access

Rate-Assured GPU Inference for 5G AI Slices

Aug 2026 · Proceedings of the ACM SIGCOMM 2026 Conference · 0 citations · 19 references

Abstract

Edge AI inference is an important workload in 5G networks. Multiple classes of edge AI workloads often share edge infrastructure, yet each may require distinct latency and rate guarantees. 5G network slicing supports differentiated requirements on the network side, but shared GPU inference also needs compute-side guarantees after requests reach the Multi-access Edge Computing (MEC) host. We extend 3GPP network slicing with compute-side enforcement so that slice guarantees remain effective after traffic reaches the MEC host. To realize this extension, we design a GPU scheduler that combines Hierarchical Token Bucket (HTB)-based traffic conditioning with Earliest Deadline First (EDF) scheduling. Our scheduler enforces per-class assured goodput, defined as the committed rate of latency-compliant completions for each class. The GPU scheduler identifies request classes via tags, which are assigned during GTPU encapsulation at the 5G user plane. This integration preserves overall latency guarantees across both the network and compute domains of a 5G slice for AI.

Read PDF