This paper proposes a new, to the authors' knowledge, efficient and accurate semantic segmentation network for LiDAR, called Range-FDSeg, and introduces a lightweight and dynamic upsampler, called Dysample-S+.
Abstract
LiDAR semantic segmentation is significant in applications such as autonomous driving and robot navigation, as it greatly improves scene perception and object detection. However, the existing methods face the challenges of achieving high segmentation accuracy while maintaining low computational cost and complexity. In this paper, we propose a new, to our knowledge, efficient and accurate semantic segmentation network for LiDAR, called Range-FDSeg. To reduce the risk of information compression and loss when projecting 3D point cloud data onto 2D range images, we design a multi-channel fusion interactive learning (FIL) module. This module effectively integrates multimodal channels, such as coordinates, depth, and reflectivity, for interactive learning. As a result, FIL module can reduce the noise interference inherent in individual channels and capture the underlying relationships between different physical quantities. To further improve the performance, we introduce a lightweight and dynamic upsampler, called Dysample-S+. It effectively resolves the inherent challenges of traditional sampling methods through its adaptive weighting mechanism, which dynamically adjusts to local geometric patterns and density variations in raw point clouds. Extensive evaluations on publicly available benchmark datasets, including SemanticKITTI, SemanticPOSS, and NuScenes, demonstrate that the proposed Range-FDSeg outperforms most existing state-of-the-art methods.
Results validate the effectiveness of the proposed novel 3D object detection and tracking framework, termed ECF3DMOT, in advancing 3D object detection and tracking for autonomous driving.
Xiaojuan Peng, Fei Teng, Tiankai Chen et al.· International Journal of Mac...· 0 citations
This paper replaces the Stem layer in FRNet with the proposed FD-Stem, which improves feature representation while reducing computational complexity, and introduces long-range modeling capability with limited additional parameters, enabling effective learning of both spatial and channel-wise representations.
Ya-Dong Guo, Jing Liu, Wei Zheng et al.· Journal of Real-Time Image P...· 0 citations
Multimodal 3D object detection is fundamental to robust perception in autonomous driving because it integrates complementary information from LiDAR and camera sensors. However, existing methods often fail to maintain robustness under out-of-distribution (OOD) corruptions caused by sensor noise, adverse weather, and env...
Zi-Ying Song, Lin Liu, Hong-Yu Pan et al.· 0 citations
This work proposes a cascade optimization framework that systematically enhances feature representation and refines multimodal fusion, and introduces the Multi-Scale Contextual Fusion Module (MSCF) to reduce alignment bias.
Semantic segmentation of large-scale 3D point clouds is a fundamental task in robotic perception, semantic mapping, and urban scene understanding. Existing methods mainly rely on geometric information, which limits their ability to distinguish semantic categories with similar spatial structures. To address this issue,...
A spatial-channel collaborative modeling framework named SPMixNet is introduced to improve fine-grained 3D scene understanding and effectively mitigates the structural and representational limitations of existing projection-based methods, providing a promising solution for accurate segmentation of small and structurall...
Hong-Dou He, Chenggong Sun, Yi-Fang Huang et al.· The Visual Computer· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.