Skip to content

LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

Extending LILA to dynamically allocate sparsity budgets via KS-scores yields state-of-the-art generative preservation at moderate compression, while uncovering fundamental single-layer architectural bottlenecks at higher compression regimes.

Abstract

Structured pruning of large language models (LLMs) offers hardware-efficient compression, yet existing methods require calibration data, gradient computation, or large auxiliary policy networks at pruning time. LILA (\emph{Latent-Informed Layer Analysis}) scores neuron importance via the Kolmogorov--Smirnov (KS) distance between empirical singular value distributions of the full and neuron-ablated feed-forward network (FFN) weight matrix, providing a closed-form spectral rule requiring no training, calibration data, or auxiliary network. Without any fine-tuning, LILA surpasses PruneNet (45M-parameter RL policy) by 1.57~pp in zero-shot accuracy on LLaMA-2-7B at 25\% sparsity, and outperforms WikiText-2-calibrated SliceGPT by up to 6.0~pp across all sparsity levels, while preserving the original architecture. After one epoch of LoRA recovery fine-tuning, LILA achieves highly competitive performance, matching the heavily calibrated SliceGPT baseline to within a 0.48~pp margin across LLaMA-2-7B and Phi-2, despite using zero calibration data. A Neural Tangent Kernel analysis confirms a 22$\times$ reduction in functional distortion versus random pruning, providing theoretical grounding for the spectral importance criterion. Finally, extending LILA to dynamically allocate sparsity budgets via KS-scores yields state-of-the-art generative preservation at moderate compression, while uncovering fundamental single-layer architectural bottlenecks at higher compression regimes.

View source

Similar papers

Preprint Aug 2026

Reservoir of Importance: Learning Semi-Structured Sparsity with Differentiable Subset Sampling

Semi-structured $N$:$M$ sparsity has emerged as a practical direction for accelerating large language models (LLMs). However, existing learnable-mask approaches incur substantial parameter and memory overhead, limiting their scalability to large models and aggressive sparsity regimes. In this work, we revisit semi-stru...

H. Dinh, Xuan Duy Ta, K. Thân et al. · 0 citations
#natural language process... Preprint Sep 2026

Debias-SparseGPT: Bias-Aware Pruning for Large Language Models

Debias-SparseGPT is introduced, a post-training pruning method incorporating representational debiasing using a second-order term defined over demographically contrasting inputs that consistently reduces pruning-induced bias compared to SparseGPT while preserving model perplexity and zero-shot accuracy.

Irina Proskurina, Guillaume Metzler, Antoine Gourru et al. · 0 citations
Book Open access Sep 2026

SSQT: A Hardware-Friendly Fusion Compression Framework of Structured Sparsification and Sensitivity-Driven Quantization for Large-Scale Language Models

The results show that compression and execution layout must be co-designed to convert model-size reduction into practical parallel inference acceleration, and that compression and execution layout must be co-designed to convert model-size reduction into practical parallel inference acceleration.

Qian-Sheng Song, Guo-Lin Tang · 0 citations
#machine learning Preprint Aug 2026

Correlation-Aware Structured Pruning for Large Language Models

Structured pruning is a promising approach for reducing the substantial inference costs of Large Language Models (LLMs) while maintaining hardware efficiency. Many existing methods assess the importance of prunable units (e.g., channels or heads) in isolation, implicitly assuming that pruning errors are additive. This...

Si-Cheng Xu, Hao Shi, Wei Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks

Adapting pretrained models to downstream tasks with limited data has become a central paradigm in modern deep learning. Yet, despite its widespread practical success, how fine-tuning leverages information from pretraining remains poorly understood theoretically. We study fine-tuning from pretrained weights through the...

A.-E. Decleves, Etienne Boursier, Nicolas Flammarion · 0 citations
#artificial intelligence Preprint Sep 2026

RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding

Penn Treebank and a 10k-vocabulary stress test confirm that geodesic decoding transfers across corpora and larger |V|, while hyperbolic curvature helps selectively, and characterize boundary collapse in naive hyperbolic recurrence and show how Mobius stabilization restores trainability.

Fang Li · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.