Skip to content
Preprint

Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs

Aug 2026 · 2 citations · 55 references
Computer Science

TL;DR

This work proposes a simple, training-free methodology compatible with existing frameworks to mitigate residual errors in calibration data activations accumulate across layers during compression, causing misalignment between representations simulated at compression time and those experienced at inference.

Abstract

Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model parameter count while maintaining task-level accuracy. However, existing SOTA frameworks share two key limitations: (1) residual errors in calibration data activations accumulate across layers during compression, causing misalignment between representations simulated at compression time and those experienced at inference; (2) the assumption that layer importance distribution is preserved post-compression does not hold. Together, these two effects introduce misalignment in the compression process in relation to the deployed model. We study these effects and propose a simple, training-free methodology compatible with existing frameworks to mitigate them, comprising: (1) Layer-by-Layer Compression with Calibration Correction; (2) Iterative Compression with Rank Allocation Correction. Implemented atop an existing SOTA decomposition framework, and evaluated on Llama and Qwen3 models across various benchmarks and compression rates, our approach demonstrates up to ~1-2.5 accuracy point improvements over per-weight and joint decomposition baselines on zero-shot tasks.

View source

Similar papers

Preprint Aug 2026

Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression

Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) are typically applied in isolation, and each hits an accuracy-efficiency wall. This thesis proposes the"Compression Trini...

Mohammad Mozaffari · 1 citation
#small language model Preprint Aug 2026

SHIFT-LLM: Distribution Shift Correction in Depth-Pruned LLMs

SHIFT-LLM, a training-free post-pruning correction framework that inserts a Linear Residual Adapter at each pruning site, consistently recovers accuracy lost to depth pruning across most configurations, achieving gains up to +15.7 points on Llama-3.1-8B-Instruct.

Ali Bahri, Hang Li, Hongliang Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Mind the Approximation: Fisher-Weighted SVD Compression for ViTs

FACTS, a structured Fisher Approximation tailored to Compressing ViTs with Fisher-weighted SVD, which enforces token-local aggregation while preserving within-token activation-gradient dependence, and introduces a fast Constrained Rank Search (CoRS), that optimizes layer-wise rank allocation while adhering to a fixed f...

Moritz Thoma, Maximilian Groezinger, Maximilian Forstenhäusler et al. · 0 citations
Preprint Aug 2026

QuaSAR: Quantization Compensation via Stable Activation-Aware Rank Truncation

This paper proposes a parameter-free truncated pseudoinverse solver which removes collapsed directions prior to inversion, and achieves 81.42\% top-1 accuracy, outperforming prior post-training methods and fine-tuning-based baselines.

Lin-Fa Lee, Yi-Yu Chang, Kuo-Hei Yeh · 0 citations
#artificial intelligence Preprint Sep 2026

On the Interaction Between Model Compression and Test-Time Adaptation

Although compressed models retain high accuracy under supervised adaptation, their TTA performance degrades significantly with increasing compression, highlighting the need to design compression strategies that preserve adaptability.

Francesco Corti, Dong Wang, Young D. Kwon et al. · 0 citations
#natural language process... Preprint Aug 2026

QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction

QUASAR is introduced, a QAT method that continuously performs lightweight, loss-aware reconstruction in the training loop to lower the loss floor and improve the resulting low-bit model, establishing QUASAR's objective as a principled optimization target.

Vincent Counathe, Ben Athiwaratkun, C. De Sa et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.