Skip to content
Preprint

Distributed Linear Programming on GPU Clusters at Extreme Scale

Sep 2026 · 0 citations · 23 references
Mathematics Computer Science

TL;DR

SHARDLP is presented, a distributed GPU LP solver that keeps the matrix and primal-dual state partitioned from sharded input through solution output and reaches the published criterion on nine of eleven instances, on the Google PDLP benchmark.

Abstract

Large linear programs can exceed the memory of a single compute node. Although first-order methods replace sparse factorizations with GPU-suited matrix-vector products, other solver phases can reintroduce a single-node memory limit. We present SHARDLP, a distributed GPU LP solver that keeps the matrix and primal-dual state partitioned from sharded input through solution output. On the Google PDLP benchmark, SHARDLP reaches the published criterion on nine of eleven instances, compared with eight in the published CPU PDLP study. On the largest benchmark, eight H200 GPUs solve a 1.185-billion-variable, 6.338-billion-nonzero LP in 9.9 minutes; the published CPU experiment reports 21.06 hours on different hardware. Beyond this benchmark, separately checked multi-node solves reach up to 13.604 billion variables and 40.807 billion nonzeros, while validated executions span up to 76 GPUs across 29 compute nodes. For column-partitioned solves, support-aware communication skips GPUs that store no coefficients for a row; on an LP with 2.76 billion nonzeros, it cuts modelled communication by 92.97% and improves solver time by 1.27x-1.52x

View source

Similar papers

#machine learning Preprint Sep 2026

GPU-Enabled Large-Scale Optimization Using Randomized Linear Algebra

This paper introduces rlaopt, a PyTorch-based package for large-scale optimization and scientific computing using randomized numerical linear algebra (RandNLA). Despite substantial progress in RandNLA-based algorithms, few implementations combine GPU acceleration with a simple interface for specifying optimization prob...

Pratik Rathore, Zachary Frangella, P. Nobel et al. · 0 citations
Preprint Sep 2026

Accelerating Branch MPC with Two-Level Parallel Direct Solves on GPUs

Branch model predictive control optimizes multiple future trajectories coupled through shared decisions, with computational demands increasing as the number of scenarios and prediction horizon grow. We present a GPU-accelerated direct linear solver for branch MPC formulations in which all trajectories share a single ro...

Feng-Long Song, Lu-Yao Zhang, Liang Wu et al. · 0 citations
#machine learning Preprint Sep 2026

GEM-KMeans: Memory-Efficient and Accurate Clustering on Massive Scale with GPU Optimization

This paper introduces GEM-KMeans, a spectrally normalized yet mathematically equivalent NLR formulation that fuses the gradient update, nonnegative projection, and sufficient statistics for normalization and iterate movement into a matrix-multiplication epilogue.

Peng Xu, Nihar Koganti, Volodymyr Kindratenko et al. · 0 citations
Book Open access Sep 2026

BAG: Faster Matrix Multiplication on a Single GPU

BAG (Basis Alternative Matrix Multiplication on GPUs), a GPU-oriented implementation of ABMM for the NVIDIA Ampere architecture is presented, designed to shrink workspace and eliminate redundant global-memory traffic, and introduce a cost-model-based recursion policy together with a Roofline-guided blocking strategy to...

Yao Liu, Ye-Wen Li, Zhong-Hai Zhang et al. · 0 citations
Preprint Sep 2026

Joint Effects of GPU Server Topology, Parallelism, and Congestion Control on MoE Inference: A Controlled Simulation Study

Mixture-of-experts (MoE) models expand capacity via sparse activation, but inference across GPUs introduces tensor-parallel (TP) collectives and expert-parallel (EP) dispatch and combine operations. Completion time depends not just on communication volume but on how logical groups map onto intra-server interconnects, G...

Kai-Kai Yuan, Rui Xi, Yu Liu · 0 citations
#artificial intelligence Preprint Sep 2026

mKernel: Fast Multi-GPU, Multi-Node Fused Kernels

Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granularity of kernels, on separate streams, reduces only part of this communication cost. Fused kernels often have better performance by transmitting each output tile as soon a...

Zi-Ming Mao, Yi-Han Zhang, S. W. Chew et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.