Aug 2026· Neural Networks· Vol 205 Pt B, pp.
109446
· 0 citations· 40 references
Medicine
TL;DR
The Multi-Scale Vision Transformer (MSViT), which integrates Spectral-Constrained Convolution for adaptive frequency-weighted patch embedding, Horizontal-Vertical Separable Attention to enforce a full-span cross-shaped effective receptive field, and Reparameterized Convolutional Position Embedding to provide boundary-aware spatial bias without the need for interpolation is proposed.
Abstract
The "train low, deploy high" paradigm offers significant practical advantages by minimizing training overhead while enabling high-fidelity inference through increased spatial resolutions. However, Vision Transformers (ViTs) often suffer from poor zero-shot generalization to unseen resolutions compared to their convolutional counterparts. We attribute this deficiency to two fundamental phenomena: intra-patch spectral drift, where image resizing suppresses discriminative mid-to-high frequency components due to interpolation-induced low-pass filtering, and inter-patch positional awareness collapse, where the interpolation of absolute position embeddings distorts spatial priors and causes the effective receptive field to degenerate into isolated patches at larger scales. To mitigate these issues, we propose the Multi-Scale Vision Transformer (MSViT), which integrates Spectral-Constrained Convolution for adaptive frequency-weighted patch embedding, Horizontal-Vertical Separable Attention to enforce a full-span cross-shaped effective receptive field, and Reparameterized Convolutional Position Embedding to provide boundary-aware spatial bias without the need for interpolation. When trained exclusively at 224 × 224, MSViT demonstrates remarkable robustness across a broad range of test resolutions, maintaining consistent and stable accuracy as the input scales from 128 × 128 up to 640 × 640. Our work underscores that explicit modeling of spectral stability and spatial structure is essential for developing resolution-flexible vision transformers.
Post-training quantization (PTQ) efficiently compresses Vision Transformers (ViTs) without retraining, yet suffers severe accuracy degradation at low bit-widths. Existing optimization-based PTQ methods guide block reconstruction via either soft logits or second-order Hessian proxies. Logit supervision is prone to overf...
Pei Sun, Guang-Qi Liang, Jin-Nian Tong et al.· 0 citations
Results indicate that the proposed framework provides an effective solution with substantially reduced boundary discontinuities for advanced neural image compression systems based on hybrid CNN-Transformer architectures.
S. Buthelezi, Jules R. Tapamo· IEEE Access· 0 citations
Aether is introduced, a simple plug-in method that applies diffusion-style random perturbations in the embedding space via controlled alpha-mixing, specifically designed to provide isotropic regularization that remains semantically consistent.
Hyesong Choi, Daeun Kim, Song Park et al.· 0 citations
Small objects in complex driving scenes often contain weak texture information and are highly sensitive to localization errors, which poses significant challenges to feature representation and bounding box regression. To address these issues, this paper proposes Acuity-DETR, a high-acuity visual perception model for sm...
The proposed framework provides a robust, scalable solution for real-world illumination enhancement across diverse lighting conditions and consistently outperforms state-of-the-art supervised and unsupervised methods in terms of fidelity, perceptual quality, and generalization.
Yasmin Yasin, Muhammad Usman, Ibrahim Radwan et al.· 0 citations
This work proposes a pixel-grounded super-resolution (PGSR) framework that preserves LR-observed pixel evidence before VAE compression and reuses it throughout restoration and improves the realism--fidelity trade-off and produces more faithful, visually convincing results than existing latent generative SR approaches.
Yu Shi, Yu-Yao Zhang, Yu-Wing Tai· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.