Skip to content
Conference Open access

When Simplicity Wins: Bottleneck-Aware Context Modeling for Lightweight Semantic Segmentation

Aug 2026 · International Conference on Information Photonics · 0 citations · 32 references
Computer Science

TL;DR

Extensive experiments demonstrate that SiConMo achieves a state-of-the-art accuracy-efficiency trade-off among lightweight semantic segmentation models, highlighting simplicity as a powerful design principle.

Abstract

Semantic segmentation demands a careful balance between accuracy, efficiency, and scalability, which remains difficult to achieve for high-resolution imagery. Convolutional networks effectively model local patterns but struggle with long-range dependencies, whereas Vision Transformers capture global context at a high computational cost. While recent work largely focuses on encoder design, the bottleneck stage, central to contextual aggregation and information flow, has been relatively overlooked. We propose SiConMo, a lightweight yet effective framework, implemented in two variants: an RGB-only model (SiConMo) and a GME-enhanced variant (SiConMo$_\dagger$). We show that simplicity arises from a key design principle: at very low computational budgets, the bottleneck is the most efficient stage to integrate local and global context. SiConMo integrates three complementary components: a Token Pyramid Extraction Module for hierarchical multi-scale representation, a Transformer-Branched Depthwise Convolution block for bottleneck-aware context modeling, and a Feature Merging Module that preserves spatial structure while enhancing semantic consistency. Extensive experiments on ADE20K, PASCAL Context, Cityscapes, and COCO-Stuff demonstrate that SiConMo achieves a state-of-the-art accuracy-efficiency trade-off among lightweight semantic segmentation models, highlighting simplicity as a powerful design principle.

Read PDF

Similar papers

Jul 2026

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

This work introduces \method, a Multi-scale Adaptive Vision Encoder, a Multi-scale Adaptive Vision Encoder that uses position-dependent gates to fuse shallow, intermediate, and deep features from a vision Transformer, preserving global semantics while enhancing edges, text, and local structure.

Sha Lei · 0 citations
Open access Aug 2026

Shared multi-branch framework and enhanced feature fusion for few-shot semantic segmentation

Traditional semantic segmentation relies on massive annotated datasets, which are often prohibitively expensive in specialized fields such as medical or satellite imagery. This paper proposes an enhanced feature extraction framework for self-support few-shot semantic segmentation that overcomes the challenge of data sc...

Jin-Ming Guo, Li-Hsuan Chen, Yi-Chong Zeng et al. · 0 citations
Open access 2026

BFANet: Fine-Grained Boundary-Aware Semantic Segmentation Driven by Dynamic Feature Alignment

: Lightweight semantic segmentation remains challenging because compact backbones often weaken feature discriminability and lose fine-grained boundary details. In DeepLabV3 + -style encoder-decoder architectures, the direct fusion of high-level semantic features and low-level spatial features may introduce semantic-spa...

Wang Zhang, Lanlan Li, Jiayi Xing et al. · 0 citations
Open access Aug 2026

Semantic Topological Multi-Scale Part Network for Fine-Grained Visual Classification

Fine-grained visual classification (FGVC) aims to distinguish highly similar subcategories, and its performance relies heavily on the accurate modeling of discriminative local parts and their structural relationships. However, existing Vision Transformer-based methods are susceptible to background noise interference, a...

Xue-Rong Liu, Min Zhi, Yan-Jun Yin et al. · 0 citations
Preprint Sep 2026

CoordFormer: Give Me Any Coordinates and I Will Give You Labels

Semantic segmentation on very-high-resolution images remains challenging due to the high computational cost and the difficulty of capturing fine-grained details. We propose CoordFormer, a novel coordinate-based architecture for semantic segmentation that predicts labels at arbitrary spatial locations through a Coordina...

Iacopo Curti, Pierluigi Zama Ramirez, Alioscia Petrelli et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.