A typed domain-specific language that captures recurring tensor structures, such as repeated regions and floating-point fields, through a set of reversible operators, is designed, which formulates lossless tensor compression as program synthesis.
Abstract
Model checkpoints are growing in both number and size, which makes archival, transfer, and deployment increasingly costly. General-purpose compressors can reduce storage requirements but ignore tensor structure, whereas existing tensor-specific compressors rely on fixed and format-specific pipelines. We present Brevis, which formulates lossless tensor compression as program synthesis. We design a typed domain-specific language (DSL) that captures recurring tensor structures, such as repeated regions and floating-point fields, through a set of reversible operators. Given a tensor, Brevis synthesizes a self-contained DSL program that reconstructs it bit-exactly. A checkpoint-specific production prior, learned from a small representative sample of tensors, guides a bounded A* search to synthesize compact programs, which can later be executed directly for bit-exact decompression. On 10 public checkpoints spanning language, audio, and image generation models, Brevis reduces 2.13 TB of checkpoint data to 1.41 TB, a 33.93% storage reduction. It produces archives up to 30.87% smaller than those of four general-purpose compressors, including zstd and gzip, and smaller archives than the tensor-specific compressors ZipNN and DFloat11. Under a practical concurrency configuration, Brevis achieves 3.60 GB/s compression and 6.61 GB/s decompression while preserving every source byte.
Modern model hubs store hundreds of petabytes of large language models (LLMs), with fine-tuned variants dominating the storage footprint. These variants contain substantial cross-model redundancy that delta compression can exploit by storing only the difference between a target and a reference model. However, compressi...
Ting-Feng Lan, Zirui Wang, Yun-Jia Zheng et al.· Proceedings of the ACM SIGOP...· 0 citations
Group-relative policy optimization (GRPO) makes learner-side activations a major memory-computation bottleneck: gradient checkpointing reduces activation memory through recomputation, but fixed schedules can leave roughly 18 GB unused on a 48-GB GPU despite substantial recomputation overhead. Existing activation-manage...
Xin-Rui Chen, Mengyang Li, Ou Wu et al.· 0 citations
Soft context compression condenses a context into a few memory tokens that a frozen LLM consumes in place of the raw text, but existing compressors fix the compression ratio at training and inference: each deployed ratio requires a separately trained model, and the chosen ratio is applied uniformly to all inputs, whose...
Kai-Yan Zhao, Zhong-Tao Miao, Akiko Aizawa et al.· 1 citation
This work introduces Schur Replay, a scale-selection algorithm that reproduces the GPTQ updates caused by each block scale and scores the resulting block error after accounting for compensation from unquantized columns.
Rui-Ying Ding, Jie Li, Kang He et al.· 0 citations
Efficient GPU implementations of tensor programs often require joint optimization of high-level algebraic formulations and low-level execution strategies. However, the resulting search space grows rapidly as transformations combine across operators, making joint optimization difficult to scale. We present EqiForge, a t...
Production machine learning (ML) stacks often split graph compilation and kernel execution across different layers and languages, making backend behavior, deployment guarantees, and performance fallbacks hard to reason about end-to-end. RLX addresses this gap with a single Rust codebase that combines compiler and runti...
Eugene Hauptmann, Nataliya Kosmyna· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.