Sep 2026· Proceedings of the International Conference on Parallel Processing· 0 citations· 33 references
TL;DR
MIGServe treats the physical layout of MIG instances as a first-class scheduling dimension through three techniques: buddy-aware partition placement, which preserves large contiguous free blocks by allocating next to existing occupied buddies; proactive pair-matching migration, which consolidates fragmented half-full buddy pairs off the critical path of inference.
Abstract
NVIDIA’s Multi-Instance GPU (MIG) technology partitions a single GPU into hardware-isolated instances of varying sizes, offering a promising substrate for serving Large Language Model (LLM) workloads with dynamic request profiles. However, existing MIG management approaches suffer from high resource fragmentation and reconfiguration overhead. The root cause is that they treat GPU resources as scalar capacities, while MIG enforces a buddy-aligned physical layout in which instances must occupy contiguous slices and can only be reshaped along fixed boundaries. This mismatch raises three challenges: (i) a layout-oblivious small instance acts as a roadblock that prevents adjacent free blocks from coalescing; (ii) stochastic request lifetimes scatter instances across the layout, accumulating fragmentation; and (iii) under skewed traffic, evicting a temporarily idle hot instance for a sporadic cold request triggers an evict-then-reload cycle. We present MIGServe, a layout-aware MIG resource management system that enables fine-grained dynamic reconfiguration for LLM serving. MIGServe treats the physical layout of MIG instances as a first-class scheduling dimension through three techniques: (1) buddy-aware partition placement, which preserves large contiguous free blocks by allocating next to existing occupied buddies; (2) proactive pair-matching migration, which consolidates fragmented half-full buddy pairs off the critical path of inference; and (3) temperature-guided eviction, which shields hot instances from transient cold requests to suppress reconfiguration thrashing. On NVIDIA A100 GPUs with production-inspired LLM workloads, MIGServe serves 2.38 × –5.32 × more requests under 90% SLO attainment, reduces fragmentation by 52.9%–100%, and cuts reconfiguration overhead by 81.7%–100% over state-of-the-art methods.
AlltoAllv communication is a critical primitive in distributed large-model inference, particularly for mixture-of-experts (MoE) models. The growing adoption of PCIe GPU systems for cost-efficient inference makes AlltoAllv performance on these systems increasingly important. Without a dedicated scale-up interconnect (e....
Yao Fei, Jin Fang, Si-Ze Zheng et al.· 0 citations
ElastiCo is presented, an elastic co-location framework that enables training and inference workloads to safely share GPUs through three integrated mechanisms that decomposes the resulting multi-resource allocation problem into per-job configuration selection subproblems via dynamic per-resource shadow prices.
Jing-Hao Wang, Yi-Hang Zhou, Xiaoyang Sun et al.· 0 citations
Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure. We present FreeToken, an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform. FreeToken co-designs the full serving stack,...
Shuo Yang, Xiao-yun Fan, Melissa Z. Pan et al.· 2 citations
Online image generation with Diffusion Transformers (DiTs) must meet latency service-level objectives (SLOs) while using GPU resources efficiently. Existing systems improve GPU utilization by batching multiple requests for joint execution. However, request-level batching offers limited control over batch size: batches...
Zhexiang Zhang, Min-Chen Yu, Yi-Fan Sun et al.· 0 citations
Weave is presented, to the authors' knowledge the first MoE overlap system that performs fine-grained dynamic SM scheduling - deciding per layer and per GPU by routing results at runtime, and achieves a 2.89x geometric-mean MoE-layer speedup and a 1.33x geometric-mean end-to-end speedup over five state-of-the-art basel...
Ziyu Huang, Yangjie Zhou, Chen-Hao Zhu et al.· 0 citations
InplaceKVCache is proposed, the first KVCache abstraction whose format fixes each byte's physical residency at write time, so that the CPU--GPU load balance can be adjusted without moving data after placement, turning load balancing into pure scheduling.
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.