Bonsai: Efficient and Optimal Automatic Tensor Rematerialization for Memory-Constrained DNN Training
GPU memory is increasingly the primary bottleneck in scaling deep neural network (DNN) training, where the activation tensors footprint of a model may exceed the memory capacity. Tensor recomputation is a powerful technique that trades additional computation for reduced peak memory usage. However, existing approaches f...