Skip to content

LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models

Jul 2026 · arXiv.org · Vol abs/2607.06918 · 0 citations · 50 references
Computer Science

TL;DR

Low-Rank Convolutional Adaptation (LoCA), a convolution-aware PEFT framework that addresses spatial-channel entanglement by decoupling channel and spatial adaptation and achieves competitive or state-of-the-art performance across fine-grained classification, domain-generalized semantic segmentation, and generative benchmarks.

Abstract

Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for diverse downstream tasks. The key challenge of VFM adaptation stems from the prohibitive costs of full fine-tuning and catastrophic forgetting. To address this, Low-Rank Adaptation (LoRA) has emerged as the prevailing paradigm for Parameter-Efficient Fine-Tuning (PEFT). However, LoRA is typically designed for transformer self-attention layers parameterized by 2D matrices. Since convolutional kernels inherently couple spatial and channel information within a 4D tensor, forcing them into a monolithic 2D matrix disrupts the inherent spatial topology. In this paper, we propose Low-Rank Convolutional Adaptation (LoCA), a convolution-aware PEFT framework that addresses spatial-channel entanglement by decoupling channel and spatial adaptation. LoCA introduces a low-rank channel adaptation for dense cross-channel mixing and refines spatial bases extracted from pre-trained kernels via Singular Value Decomposition (SVD). Experimental results show that LoCA preserves pre-trained spatial priors and achieves competitive or state-of-the-art performance across fine-grained classification, domain-generalized semantic segmentation, and generative benchmarks.

View source

Similar papers

Aug 2026

Extending the scale generalization of the Vision Transformer without fine-tuning.

The Multi-Scale Vision Transformer (MSViT), which integrates Spectral-Constrained Convolution for adaptive frequency-weighted patch embedding, Horizontal-Vertical Separable Attention to enforce a full-span cross-shaped effective receptive field, and Reparameterized Convolutional Position Embedding to provide boundary-a...

Kai Jiang, Peng Peng, Youzao Lian et al. · 0 citations
Preprint Aug 2026

TASSO: TAsk-Specific Subspace Optimization for Continual Learning of Vision-Language Models

TASSO, a new paradigm that efficiently preserves the latent space geometry while ensuring network plasticity, is introduced with two complementary techniques: subspace learning and geometry-aware knowledge distillation.

Chang-Ming Sun, Francesco Barbato, Matteo Caligiuri et al. · 0 citations
Aug 2026

SPIRA: Sparse Information-Geometric Rank Adaptation for Parameter-Efficient Fine-Tuning of Large Pretrained Models.

Downstream adaptation of large pretrained models (LPMs) via full-parameter fine-tuning is computationally prohibitive. Parameter-efficient fine-tuning (PEFT) methods, such as the widely used Low-Rank Adaptation (LoRA), reduce this cost but still parameterize dense updates over the selected weight matrices. This support...

Zhongyi Wen, Zhikai Zhai, Guo-Min Sun et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TaRA: Training-Aware Low-Rank Adaptation Initialization

Low-Rank Adaptation (LoRA) has become a de facto standard for parameter-efficient fine-tuning (PEFT), yet its performance is highly sensitive to initialization due to the information bottleneck imposed by low-rank decomposition. Existing approaches attempt to construct high-quality LoRA initializations by exploiting pr...

Taehyeon Kim, Eunhyeok Park · 0 citations
Preprint Aug 2026

P3CA: Encoder-Agnostic Interpretation of Vision Foundation Model Embeddings via Spatial Probing

Position-prompted PCA (P3CA), an encoder-agnostic method for local probing of channel-rich spatial tensors, is proposed and implemented in EmbedVision, an interactive 3D Slicer-based workflow, and evaluated across natural images, colorectal pathology foundation-model embeddings, and spatial transcriptomic tensors.

A. Jamzad, Dilakshan Srikanthan, F. Akbarifar et al. · 0 citations
Aug 2026

EP-MAE: A resource-efficient masked autoencoding framework for 3D neural representation learning.

Efficient Point Masked Autoencoders (EP-MAE), a new framework designed to significantly reduce the training cost of 3D self-supervised pre-training while maintaining strong representation quality, and provides a scalable and effective foundation for future 3D neural network models is presented.

Jian Zhu, Jiale Zhao, Cheng Lin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.