Aug 2026· International Conference on Information Photonics· 0 citations· 32 references
Computer Science
TL;DR
Extensive experiments demonstrate that SiConMo achieves a state-of-the-art accuracy-efficiency trade-off among lightweight semantic segmentation models, highlighting simplicity as a powerful design principle.
Abstract
Semantic segmentation demands a careful balance between accuracy, efficiency, and scalability, which remains difficult to achieve for high-resolution imagery. Convolutional networks effectively model local patterns but struggle with long-range dependencies, whereas Vision Transformers capture global context at a high computational cost. While recent work largely focuses on encoder design, the bottleneck stage, central to contextual aggregation and information flow, has been relatively overlooked. We propose SiConMo, a lightweight yet effective framework, implemented in two variants: an RGB-only model (SiConMo) and a GME-enhanced variant (SiConMo$_\dagger$). We show that simplicity arises from a key design principle: at very low computational budgets, the bottleneck is the most efficient stage to integrate local and global context. SiConMo integrates three complementary components: a Token Pyramid Extraction Module for hierarchical multi-scale representation, a Transformer-Branched Depthwise Convolution block for bottleneck-aware context modeling, and a Feature Merging Module that preserves spatial structure while enhancing semantic consistency. Extensive experiments on ADE20K, PASCAL Context, Cityscapes, and COCO-Stuff demonstrate that SiConMo achieves a state-of-the-art accuracy-efficiency trade-off among lightweight semantic segmentation models, highlighting simplicity as a powerful design principle.
This work introduces \method, a Multi-scale Adaptive Vision Encoder, a Multi-scale Adaptive Vision Encoder that uses position-dependent gates to fuse shallow, intermediate, and deep features from a vision Transformer, preserving global semantics while enhancing edges, text, and local structure.
Traditional semantic segmentation relies on massive annotated datasets, which are often prohibitively expensive in specialized fields such as medical or satellite imagery. This paper proposes an enhanced feature extraction framework for self-support few-shot semantic segmentation that overcomes the challenge of data sc...
Jin-Ming Guo, Li-Hsuan Chen, Yi-Chong Zeng et al.· APSIPA Transactions on Signa...· 0 citations
Results indicate that BCSNet provides a practical accuracy--efficiency trade-off for high-resolution real-time segmentation, while direct embedded deployment and hardware-specific optimization remain directions for future work.
Qingpei ​LIU· Poster Volume 0007 The 2026...· 0 citations
: Lightweight semantic segmentation remains challenging because compact backbones often weaken feature discriminability and lose fine-grained boundary details. In DeepLabV3 + -style encoder-decoder architectures, the direct fusion of high-level semantic features and low-level spatial features may introduce semantic-spa...
Wang Zhang, Lanlan Li, Jiayi Xing et al.· Computers, Materials & C...· 0 citations
Fine-grained visual classification (FGVC) aims to distinguish highly similar subcategories, and its performance relies heavily on the accurate modeling of discriminative local parts and their structural relationships. However, existing Vision Transformer-based methods are susceptible to background noise interference, a...
Xue-Rong Liu, Min Zhi, Yan-Jun Yin et al.· Journal of Imaging· 0 citations
Semantic segmentation on very-high-resolution images remains challenging due to the high computational cost and the difficulty of capturing fine-grained details. We propose CoordFormer, a novel coordinate-based architecture for semantic segmentation that predicts labels at arbitrary spatial locations through a Coordina...