This work proposes a simple, training-free methodology compatible with existing frameworks to mitigate residual errors in calibration data activations accumulate across layers during compression, causing misalignment between representations simulated at compression time and those experienced at inference.
Abstract
Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model parameter count while maintaining task-level accuracy. However, existing SOTA frameworks share two key limitations: (1) residual errors in calibration data activations accumulate across layers during compression, causing misalignment between representations simulated at compression time and those experienced at inference; (2) the assumption that layer importance distribution is preserved post-compression does not hold. Together, these two effects introduce misalignment in the compression process in relation to the deployed model. We study these effects and propose a simple, training-free methodology compatible with existing frameworks to mitigate them, comprising: (1) Layer-by-Layer Compression with Calibration Correction; (2) Iterative Compression with Rank Allocation Correction. Implemented atop an existing SOTA decomposition framework, and evaluated on Llama and Qwen3 models across various benchmarks and compression rates, our approach demonstrates up to ~1-2.5 accuracy point improvements over per-weight and joint decomposition baselines on zero-shot tasks.
Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) are typically applied in isolation, and each hits an accuracy-efficiency wall. This thesis proposes the"Compression Trini...
SHIFT-LLM, a training-free post-pruning correction framework that inserts a Linear Residual Adapter at each pruning site, consistently recovers accuracy lost to depth pruning across most configurations, achieving gains up to +15.7 points on Llama-3.1-8B-Instruct.
Ali Bahri, Hang Li, Hongliang Li et al.· 0 citations
FACTS, a structured Fisher Approximation tailored to Compressing ViTs with Fisher-weighted SVD, which enforces token-local aggregation while preserving within-token activation-gradient dependence, and introduces a fast Constrained Rank Search (CoRS), that optimizes layer-wise rank allocation while adhering to a fixed f...
Moritz Thoma, Maximilian Groezinger, Maximilian Forstenhäusler et al.· 0 citations
This paper proposes a parameter-free truncated pseudoinverse solver which removes collapsed directions prior to inversion, and achieves 81.42\% top-1 accuracy, outperforming prior post-training methods and fine-tuning-based baselines.
Although compressed models retain high accuracy under supervised adaptation, their TTA performance degrades significantly with increasing compression, highlighting the need to design compression strategies that preserve adaptability.
Francesco Corti, Dong Wang, Young D. Kwon et al.· 0 citations
QUASAR is introduced, a QAT method that continuously performs lightweight, loss-aware reconstruction in the training loop to lower the loss floor and improve the resulting low-bit model, establishing QUASAR's objective as a principled optimization target.
Vincent Counathe, Ben Athiwaratkun, C. De Sa et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.