Witcher: Training Large Language Models With Heterogeneous Spot GPUs in Data Centers
As the size of large language models (LLMs) continues to grow, training these models typically relies on data centers equipped with high-performance GPUs. However, in modern data centers, acquiring large-scale homogeneous GPU resources often results in long queueing delays, and training with on-demand instances incurs...