TensorDex: A Compact, Tensor-Centric Storage System for Modern AI Models
Abstract
Modern model hubs store hundreds of petabytes of large language models (LLMs), with fine-tuned variants dominating the storage footprint. These variants contain substantial cross-model redundancy that delta compression can exploit by storing only the difference between a target and a reference model. However, compression effectiveness depends critically on choosing a similar reference. At model-hub scale, this is challenging because model lineage metadata is often missing or unreliable, and different tensors within the same model may be most similar to tensors from different models. Consequently, model-level pairing leaves substantial redundancy unexploited. We present TensorDex, a lossless, tensor-centric storage system that performs delta compression at tensor granularity. TensorDex decomposes models into tensors, uses compact bit-level fingerprints to predict tensor-pair compressibility, and incrementally organizes tensors into multi-center clusters to select effective bases as new models arrive. Evaluated on 2,890 randomly sampled Hugging Face models, Tensor-Dex reduces storage footprint by 70.5%, 37% lower than the state-of-the-art design. It achieves 22.9 GB/s compression and 28.4 GB/s decompression throughput, 3.86× and 1.49× faster than the next-best system, respectively.