Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Book Open access Sep 2026

MIGServe: Layout-Aware Multi-Instance GPU Management for Efficient LLM Serving

MIGServe treats the physical layout of MIG instances as a first-class scheduling dimension through three techniques: buddy-aware partition placement, which preserves large contiguous free blocks by allocating next to existing occupied buddies; proactive pair-matching migration, which consolidates fragmented half-full b...

Jian-Wen Chen, Yun-Kai Liang, Bin Gao et al. · 0 citations
#large language models Book Open access Sep 2026

MIGServe: Layout-Aware Multi-Instance GPU Management for Efficient LLM Serving

MIGServe treats the physical layout of MIG instances as a first-class scheduling dimension through three techniques: buddy-aware partition placement, which preserves large contiguous free blocks by allocating next to existing occupied buddies; proactive pair-matching migration, which consolidates fragmented half-full b...

Jian-Wen Chen, Yun-Kai Liang, Bin Gao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.