Skip to content

Author

Yen-Ting Kuo

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Aug 2026

Rate-Assured GPU Inference for 5G AI Slices

Edge AI inference is an important workload in 5G networks. Multiple classes of edge AI workloads often share edge infrastructure, yet each may require distinct latency and rate guarantees. 5G network slicing supports differentiated requirements on the network side, but shared GPU inference also needs compute-side guarantees after requests reach the Multi-access Edge Computing (MEC) host. We extend 3GPP network slicing with compute-side enforcement so that slice guarantees remain effective after traffic reaches the MEC host. To realize this extension, we design a GPU scheduler that combines Hierarchical Token Bucket (HTB)-based traffic conditioning with Earliest Deadline First (EDF) scheduling. Our scheduler enforces per-class assured goodput, defined as the committed rate of latency-compliant completions for each class. The GPU scheduler identifies request classes via tags, which are assigned during GTPU encapsulation at the 5G user plane. This integration preserves overall latency guarantees across both the network and compute domains of a 5G slice for AI.

Yu-Hong Shen, Wen-Ju Chiang, Hsiang-Ming Hung et al. · 0 citations