Skip to content

Broadcast-and-Mixing Transformer for 3D Semantic Segmentation

Aug 2026 · IEEE Transactions on Image Processing · Vol 35, pp. 9226-9240 · 0 citations · 106 references
Medicine

Abstract

Transformers have shown great promise in various point cloud comprehension tasks, but still face challenges due to the quadratic computational and memory cost when dealing with large-scale 3D point clouds. Many recent studies focus on reducing these costs and improving model performance by solely applying restricted local attention but overlook the coarse-grained global structural information, which is also crucial to 3D semantic segmentation. In this paper, we propose a novel Broadcast-and-Mixing Transformer model for 3D semantic segmentation. Leveraging the joint utilization of global, regional, and local structures within the point cloud, our approach first broadcasts the global representations learned by a lightweight voxel set attention to the regional level and then mixes them with local point features using a unique voxel–point self-attention mechanism. The model enables effective information exchange across different granularity levels, encompassing global-regional-local interactions, and controlling the overall computational complexity without a substantial increase after incorporating global information. Extensive experiments on large-scale indoor and outdoor datasets demonstrate the effectiveness of our proposed method, surpassing hybrid-input approaches and matching global-attention baselines with significantly lower memory cost.

View source

Similar papers

Conference Open access Sep 2026

PointGP: Geometry-Primed Attention for Point Cloud Analysis

PointGP is proposed, a geometry-primed framework that uses rectified local geometric topology as the primary cue for attention generation and achieves competitive accuracy with strong parameter and computational efficiency compared with representative strong baselines.

Yong Yang, Jian-Min Huang, Meng-Yuan Ge et al. · 0 citations
Open access Sep 2026

Vision–language guided semantic-geometric transformer for memory-efficient 3D scene understanding

This work proposes a segmentation framework guided by Contrastive Language–Image Pre-training (CLIP) that enriches sparse 3D tokens with vision–language semantic priors and introduces a decoupled CLIP-induced semantic residual that forms semantic-geometric attention biases for local window attention.

Li-Cheng Liu, Yu Li, Fu-Yong Liu · 0 citations
2026

Coupled Geometry-Sequence State-Space Model for Point Cloud Semantic Segmentation

The complex spatial structures and varying scales of point cloud data make semantic segmentation a highly challenging task. While downsampling and feature fusion are essential steps in mainstream networks, existing models often suffer from local geometry loss and multiscale semantic conflicts during these processes. To...

Dong-Rong Yang, Si-Yuan Hao, Yuan-Xin Ye · 0 citations
Preprint Aug 2026

Stream3Dv2: Geometric-Semantic Fusion Enhanced Streaming Zero-Shot 3D Scene Understanding

A comprehensive geometric-semantic fusion mechanism that resolves geometric noise and semantic ambiguity by explicitly utilizing semantic guidance and formulating 3D segmentation as solving point-and-set merging and partitioning problems, and an innovative manifold-distance-based point cloud refinement strategy.

Jie Xu, Na Zhao · 0 citations
Preprint Sep 2026

Partition-Invariant Tuning for 3D Scene Understanding

PointPiT is proposed, a partition-invariant tuning framework for scene-level point clouds that integrates local geometric patterns with global scene context to mitigate partition-induced representation shifts, and achieves consistent state-of-the-art performance among representative PEFT methods.

Hong-Qiang Lin, Tian-Le Wang, Shui-Wang Li et al. · 0 citations
Conference Open access Sep 2026

Interactive Open-Set Semantic Mapping with a 3D Scene Graph Backend

A modular mapping architecture is demonstrated that establishes 3D Semantic Scene Graphs (3DSSGs) as its foundational back-end, enabling the dense representation of extensive environments containing thousands of unique object instances and supporting open-vocabulary queries via CLIP features without requiring any addit...

Felix Igelbrink, Lennart Niecksch, Martin Günther et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.