Witcher: Training Large Language Models With Heterogeneous Spot GPUs in Data Centers
Abstract
As the size of large language models (LLMs) continues to grow, training these models typically relies on data centers equipped with high-performance GPUs. However, in modern data centers, acquiring large-scale homogeneous GPU resources often results in long queueing delays, and training with on-demand instances incurs prohibitively high financial costs, limiting the efficiency of large-scale model training. In this paper, we present Witcher, a system for efficiently training large language models with heterogeneous spot GPUs in data centers. It utilizes more easily obtained heterogeneous GPUs and low-cost, preemptible spot instances, while still achieving stable and efficient training performance. Witcher introduces a heterogeneous training framework that supports asymmetric 3D parallelism, allowing flexible and efficient parallel configurations across heterogeneous GPUs. To address frequent preemption and allocation events of spot instances, Witcher employs a fast recovery mechanism that reconstructs interrupted pipelines by copying from intact model replicas. Witcher further develops an efficient automatic parallelism algorithm that includes two dynamic programming approaches, pipeline construction and micro-batch distribution, which minimize the training iteration time. Evaluation on multiple LLMs using two real-world traces of heterogeneous spot GPUs shows that Witcher achieves up to $14.8\times $ higher training throughput compared to state-of-the-art frameworks.