Skip to content

Mind the Approximation: Fisher-Weighted SVD Compression for ViTs

Sep 2026 · 0 citations · 60 references
Computer Science

TL;DR

FACTS, a structured Fisher Approximation tailored to Compressing ViTs with Fisher-weighted SVD, which enforces token-local aggregation while preserving within-token activation-gradient dependence, and introduces a fast Constrained Rank Search (CoRS), that optimizes layer-wise rank allocation while adhering to a fixed floating point operation (FLOP) constraint.

Abstract

Model compression is key to mitigate deployment challenges of ever growing machine learning models. In this area of research, singular value decomposition (SVD)-based compression offers a compelling trade-off between computational efficiency and model accuracy. Fisher-weighted SVD in particular provides principled, loss-aware compression. However, we find that improving the fidelity of Fisher approximation used in the compression is poorly predictive of post-compression accuracy for Vision Transformers (ViTs). Motivated by this observation, we propose FACTS, a structured Fisher Approximation tailored to Compressing ViTs with Fisher-weighted SVD, which enforces token-local aggregation while preserving within-token activation-gradient dependence. Additionally, we introduce a fast Constrained Rank Search (CoRS), that optimizes layer-wise rank allocation while adhering to a fixed floating point operation (FLOP) constraint. Extensive experiments across ViTs and hybrid architectures demonstrate that FACTS consistently improves accuracy-efficiency trade-offs without requiring finetuning. Notably, it outperforms the strongest SVD baseline by up to +5.8 percentage points (p.p.) Top-1 on Swin-B, with further gains driven by our search method. Code is available at https://github.com/MoritzTho/FACTS.

View source

Similar papers

#machine learning Preprint Sep 2026

Learning Functional Subspaces for Neural Network Compression

Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keeping the matrices dense, and thus efficient on standard hardware. Existing methods, however, choose the subspace to remove from each weight matrix with local closed-form crit...

Massimo Bini, Anders Christensen, S. Alaniz et al. · 0 citations

UvA-DARE (Digital Academic Repository) Elastic ViTs from Pretrained Models without Retraining

SnapViT: single-shot network approximation for pruned Vision Transformers is introduced, a new post-pretraining structured pruning method that enables elastic inference across a continuum of compute budgets, and a self-supervised importance scoring mechanism that maintains strong performance without requiring retrainin...

Walter Simoncini, Michael Dorkenwald, Tijmen Blankevoort et al. · 0 citations
Book Open access Sep 2026

SSQT: A Hardware-Friendly Fusion Compression Framework of Structured Sparsification and Sensitivity-Driven Quantization for Large-Scale Language Models

The results show that compression and execution layout must be co-designed to convert model-size reduction into practical parallel inference acceleration, and that compression and execution layout must be co-designed to convert model-size reduction into practical parallel inference acceleration.

Qian-Sheng Song, Guo-Lin Tang · 0 citations
Open access Sep 2026

A novel approach to lossless convolutional neural network compression via progressive knowledge distillation-incorporated low-rank compression.

Model compression is widely used to deploy large neural networks on resource-constrained edge devices. Among existing techniques, low-rank composition is theoretically grounded in approximation theory and provides a strong basis for preserving model performance after compression. However, in practice, even state-of-the...

Ya-Ping He, Hao Wu, Wei-Bo Liu et al. · 0 citations
#natural language process... Preprint Sep 2026

MILO: Efficient Many-shot In-Context Learning with Block-wise Low-rank Compression

Many-shot in-context learning (ICL) enables large language models (LLMs) to adapt to complex tasks by conditioning on thousands of demonstration examples, but this paradigm shifts the inference efficiency bottleneck to the key-value (KV) cache memory. Due to the linear scaling behavior of the KV cache, storing these in...

You-Peng Zhao, Tian Tan, Li-Qian Peng et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.