Skip to content

Train Overcomplete, Deploy Compact: Scaling Recovery Capacity for Structured LLM Pruning

Sep 2026 · 0 citations · 34 references
Computer Science

TL;DR

OverRep is proposed, an Overcomplete Reparameterization framework for structured LLM pruning that temporarily overparameterizes the recovery module during training to absorb complex knowledge distilled from the original model.

Abstract

Large language models achieve strong performance across diverse tasks, but deployment remains costly because of memory, latency, and energy demands. Structured pruning reduces these costs by removing architectural components, yet its recovery stage is often limited by a mismatch between the recovery module's representational capacity and the complexity of the removed knowledge. We call this bottleneck the capacity-knowledge asymmetry and propose OverRep, an Overcomplete Reparameterization framework for structured LLM pruning. Following the principle of"train overcomplete, deploy compact", OverRep temporarily overparameterizes the recovery module during training to absorb complex knowledge distilled from the original model. After recovery, the overcomplete re-parameterization is algebraically merged into a mathematically equivalent compact module, preserving the pruned model's inference-time architecture and computational cost. OverRep further introduces an annealed activation that enables nonlinear training dynamics while converging to a linear regime for exact algebraic merging. Across three backbone families, OverRep improves retained reasoning performance over strong recovery baselines by up to 5.5 and 8.4 points at 25% and 50% pruning, respectively, while keeping memory usage and TFLOPs comparable to existing recovery methods. Our code is available at https://github.com/mmai-laboratory/OverRep.

View source

Similar papers

#machine learning Preprint Aug 2026

Correlation-Aware Structured Pruning for Large Language Models

Structured pruning is a promising approach for reducing the substantial inference costs of Large Language Models (LLMs) while maintaining hardware efficiency. Many existing methods assess the importance of prunable units (e.g., channels or heads) in isolation, implicitly assuming that pruning errors are additive. This...

Si-Cheng Xu, Hao Shi, Wei Zhang et al. · 0 citations
Preprint Sep 2026

Linear Reusable Neural Bases Architecture for Network Compression

Memory constraints remain a critical bottleneck in the deployment of large-scale AI models. Parameter sharing across network depth reduces model storage, but repeatedly applying an identical transformation limits flexibility across layers. Inspired by time--memory trade-offs in classical algorithms, we introduce the Li...

Bin-Shuai Wang, Peng Wei, Mahyar Ghazanfari · 0 citations
#machine learning Preprint Sep 2026

Output-aware Residual Stream Pruning for Large Language Models

Residual stream pruning methods reduce inference cost by shrinking the model's hidden dimension, but existing approaches typically choose these dimensions by minimizing activation reconstruction error. This criterion implicitly treats all perturbation directions as equally important, ignoring the sensitivity of downstr...

Chayne Thrash, Ke Chen, Soheil Kolouri · 0 citations
#machine learning Preprint Aug 2026

Task-Aware Spectral Pruning: A Mixture-of-Masks Framework for Efficient LLM Inference

Task-Aware Spectral Pruning (TASP), a post-training framework that calibrates module-level spectral descriptors against measured task-specific ablation effects, closes grouped-query-attention and SwiGLU dependencies during sparse-mask construction, and routes each user turn to one compiled mask that remains fixed throu...

Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter et al. · 0 citations
#natural language process... Preprint Aug 2026

Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference

Reduced Matrix Multiplication is proposed, a training-free, input-adaptive inference method that reduces Transformer matrix products by selecting informative slices along their contraction dimensions, without modifying model weights, and it is shown that the same principle extends to multimodal vision-language inferenc...

Zi-Xuan Lan, Yan-Hong Li, Jia-Wei Zhou · 0 citations
Open access Sep 2026

MoEP: Compact and efficient sparsity with modular expert paths.

The transition from dense to sparse model architectures has become a key trend in the field of Large Language Models (LLMs). Mixture-of-Experts (MoE) methods can be used to increase model conditional representation capacity by activating only a subset of parameters for each input token. However, the practical efficienc...

Joonas Tapaninaho, Mourad Oussalah · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.