Skip to content

VLMSNet: View Large to Measure Shape

Sep 2026 · IEEE Transactions on Image Processing · Vol 35, pp. 10272-10285 · 0 citations
Medicine

Abstract

Recent advances in single image super-resolution (SISR) have leveraged convolutional neural networks (CNNs) and vision transformers to model pixel-level statistics, often relying on increasingly complex architectures to capture spatial correlations. However, these approaches generally overlook a fundamental distinction between human and machine perception, that is, the human visual system prioritizes shape-centric and semantically guided interpretation over raw pixel fidelity. To bridge this gap, we propose the View Large to Measure Shape Network (VLMSNet), a novel SISR framework inspired by the shape-centric strategy of the human visual system. Specifically, VLMSNet first employs a Group Mask Generator (GMG) to derive shape-aware guidance from large-receptive-field features, providing structurally informed cues for reconstruction. An Omni Attention Block (OAB) is further introduced to jointly model spatial and channel dependencies over broad contexts, so as to better represent complex structures and textures. In addition, VLMSNet adopts a dual-branch architecture with Foreground Feature Extraction (FFE) and Background Feature Extraction (BFE), explicitly separating salient object structures from less informative regions for more targeted and efficient feature learning. Extensive experiments both on standard SISR benchmarks and real-world scenarios demonstrate that VLMSNet achieves superior performance over existing state-of-the-art methods in most cases, confirming its robustness and practical applicability.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.