Skip to content

Extending the scale generalization of the Vision Transformer without fine-tuning.

Aug 2026 · Neural Networks · Vol 205 Pt B, pp. 109446 · 0 citations · 40 references
Medicine

TL;DR

The Multi-Scale Vision Transformer (MSViT), which integrates Spectral-Constrained Convolution for adaptive frequency-weighted patch embedding, Horizontal-Vertical Separable Attention to enforce a full-span cross-shaped effective receptive field, and Reparameterized Convolutional Position Embedding to provide boundary-aware spatial bias without the need for interpolation is proposed.

Abstract

The "train low, deploy high" paradigm offers significant practical advantages by minimizing training overhead while enabling high-fidelity inference through increased spatial resolutions. However, Vision Transformers (ViTs) often suffer from poor zero-shot generalization to unseen resolutions compared to their convolutional counterparts. We attribute this deficiency to two fundamental phenomena: intra-patch spectral drift, where image resizing suppresses discriminative mid-to-high frequency components due to interpolation-induced low-pass filtering, and inter-patch positional awareness collapse, where the interpolation of absolute position embeddings distorts spatial priors and causes the effective receptive field to degenerate into isolated patches at larger scales. To mitigate these issues, we propose the Multi-Scale Vision Transformer (MSViT), which integrates Spectral-Constrained Convolution for adaptive frequency-weighted patch embedding, Horizontal-Vertical Separable Attention to enforce a full-span cross-shaped effective receptive field, and Reparameterized Convolutional Position Embedding to provide boundary-aware spatial bias without the need for interpolation. When trained exclusively at 224 × 224, MSViT demonstrates remarkable robustness across a broad range of test resolutions, maintaining consistent and stable accuracy as the input scales from 128 × 128 up to 640 × 640. Our work underscores that explicit modeling of spectral stability and spatial structure is essential for developing resolution-flexible vision transformers.

View source

Similar papers

Preprint Sep 2026

GLF-Q: Global-Local Feature-based Quantization for Vision Transformers

Post-training quantization (PTQ) efficiently compresses Vision Transformers (ViTs) without retraining, yet suffers severe accuracy degradation at low bit-widths. Existing optimization-based PTQ methods guide block reconstruction via either soft logits or second-order Hessian proxies. Logit supervision is prone to overf...

Pei Sun, Guang-Qi Liang, Jin-Nian Tong et al. · 0 citations
Preprint Sep 2026

Isotropic Embedding Perturbations for Robust Vision Language Encoders

Aether is introduced, a simple plug-in method that applies diffusion-style random perturbations in the embedding space via controlled alpha-mixing, specifically designed to provide isotropic regularization that remains semantically consistent.

Hyesong Choi, Daeun Kim, Song Park et al. · 0 citations
Open access 2026

Acuity-DETR: Enhancing Perceptual Acuity for Small Object Detection in Driving Scenes

Small objects in complex driving scenes often contain weak texture information and are highly sensitive to localization errors, which poses significant challenges to feature representation and bounding box regression. To address these issues, this paper proposes Acuity-DETR, a high-acuity visual perception model for sm...

Ling-Ling Wang, Xiang Li, Shu-Ling Yin et al. · 0 citations
Preprint Aug 2026

UBLLIE: Unified Backlight and Low-Light Image Enhancement

The proposed framework provides a robust, scalable solution for real-world illumination enhancement across diverse lighting conditions and consistently outperforms state-of-the-art supervised and unsupervised methods in terms of fidelity, perceptual quality, and generalization.

Yasmin Yasin, Muhammad Usman, Ibrahim Radwan et al. · 0 citations
Preprint Aug 2026

When Latents Forget Pixels: Restoring Fidelity in Diffusion Transformer Super-Resolution

This work proposes a pixel-grounded super-resolution (PGSR) framework that preserves LR-observed pixel evidence before VAE compression and reuses it throughout restoration and improves the realism--fidelity trade-off and produces more faithful, visually convincing results than existing latent generative SR approaches.

Yu Shi, Yu-Yao Zhang, Yu-Wing Tai · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.