Skip to content
Preprint

SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs

Sep 2026 · 0 citations · 35 references
Computer Science

TL;DR

SlideDP is presented, a synchronous data-parallel runtime for shared-host multi-GPU systems that maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks.

Abstract

Host-resident layer streaming enables full-parameter LLM fine-tuning beyond GPU memory, but data-parallel ranks compete for shared host resources. Replicated transfers amplify traffic, while strong scaling can expose host work as computation windows shrink. We present SlideDP, a synchronous data-parallel runtime for shared-host multi-GPU systems. It maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks. An analytical step-time model characterizes resource bottlenecks and pipeline exposure; runtime measurements guide communication, chunking, and activation policies under a GPU memory budget. In matched-batch sweeps, SlideDP achieves geometric-mean throughput ratios of 1.46-2.64$\times$ over SlideFormer, MegaTrain, and ZeRO-Offload. On four H100s, SlideDP approaches GPU-resident FSDP2 throughput for Qwen3-14B at a smaller batch size. With a larger batch, it processes over 1M tokens per step and exceeds FSDP2's measured peak throughput by 11.2%. Separately, it supports 256K-token sequences for the same model and fine-tunes Qwen2.5-72B on four RTX 4090 GPUs. Project page: https://github.com/RegiaYoung/SlideDP.

View source

Similar papers

Preprint Aug 2026

LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training

LazyTrain is proposed, an optimization layer over a layer-streaming executor that formulates checkpoint selection, activation placement, recomputation, and CPU-GPU-NVMe communication overlap as a mixed-integer scheduling problem, then executes the solved policy during training.

Xiao-Jun Wu, Ce-Hao Yang, Hong-Hao Liu et al. · 1 citation
#artificial intelligence Preprint Sep 2026

EMA: Elastic and Performance Transparent Memory Across GPUs

Multi-GPU servers have become the standard building block of modern data centers, providing aggregated capacity through high-bandwidth interconnects. At the same time, workloads such as LLM inference exhibit highly dynamic memory demands, which can cause one GPU to exhaust its local memory while others remain underutil...

Yi Xu, Tian Xia, Ion Stoica · 0 citations
Open access Sep 2026

ZenFlow: Enabling Stall-Free Offloading for LLM Training

Fine-tuning large language models (LLMs) often exceeds GPU memory limits, prompting systems to offload model states to CPU memory. However, existing offloaded training frameworks like ZeRO-Offload treat all parameters equally and update the full model on the CPU, causing severe GPU stalls, where fast, expensive GPUs si...

Ting-Feng Lan, Yu-Sen Wu, Bin Ma et al. · 0 citations
Preprint Aug 2026

OpRAG: A Resource-Deterministic Runtime for GPU-Backed Multi-Stage RAG Workflows

OpRAG is presented, a resource-deterministic distributed runtime for GPU-backed multi-stage RAG workflows that combines an Arrow zero-copy data plane, persistent workers, bounded queues, CPU tokenizer prefetching, batched GPU embedding, and overlapped retrieval/generation execution to reduce non-model overhead around L...

A. Sarker, M. Staylor, Aymen Alsaadi et al. · 1 citation
Preprint Sep 2026

HOCCL: Offloading Collective Communication from GPU Cores to Accelerate Distributed Training

Large language model training involves massive computation on GPU streaming multiprocessors (SMs), the primary compute units of GPUs. Since SMs host specialized accelerators such as Tensor Cores, their efficient utilization is critical to training efficiency. Unfortunately, existing collective communication systems com...

Yao Fei, Gong-Ming Zhao, Hong-Li Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.