Jul 2026· Measurement science and technology· Vol 37, pp. 295404· 0 citations· 34 references
TL;DR
STRDet, a robust 3D object detection framework based on progressive spatio-temporal feature refinement, is proposed and a Semantic-guided Context Refinement (SCR) module that explicitly suppresses background interference prior to the view transformation, thereby blocking noise propagation is proposed.
Abstract
Camera-based 3D object detection has attracted widespread attention for autonomous driving applications. However, existing methods often lack effective feature screening mechanisms, resulting in an extremely low spatio-temporal signal-to-noise ratio in complex scenes. Specifically, distracting background projections, feature misalignment caused by dynamic objects, and frequent occlusions jointly lead to severe ambiguity and loss of object features. To alleviate these issues, we propose STRDet, a robust 3D object detection framework based on progressive spatio-temporal feature refinement. First, we propose a Semantic-guided Context Refinement (SCR) module that explicitly suppresses background interference prior to the view transformation, thereby blocking noise propagation. Second, we design the Differential-aware Feature Alignment (DFA) and Residual-based Adaptive Gated Fusion (RAGF) modules, which leverage feature difference maps as motion saliency indicators to guide deformable alignment and employ gating mechanisms to selectively integrate historical motion cues, effectively resolving dynamic alignment failures and mitigating feature loss under occlusion. Extensive experiments on the nuScenes dataset demonstrate that STRDet effectively enhances feature purity and coherence, yielding significant improvements and achieving 46.95\% mAP and 55.37\% nuScenes detection score.
This work introduces SAM-AD, a domain-specific pretraining strategy that fine-tunes SAM on autonomous-driving imagery to extract feature representations with rich semantic information, and develops the Depth-Guided Wavelet Attention (DGWA) module, which suppresses high-frequency sensor noise while preserving critical c...
Zi-Ying Song, Lin Liu, Hong-Yu Pan et al.· 0 citations
A depth uncertainty-guided feature modulation method is proposed, in which depth entropy and variance are jointly modeled to generate a BEV alignment confidence map, enabling adaptive enhancement and suppression of image features and effectively mitigating cross-modal alignment errors.
Jie Hu, Xinghao Cheng, Shuaidi He et al.· International Conference on...· 0 citations
This research proposes a novel framework that enhances the BEV representation with temporal modeling, and effectively enhances 3D detection accuracy in dynamic scenarios.
Meng-Jia Shao, Wei Li, Jie Bai et al.· SAE technical paper series· 0 citations
Multi-Object Tracking (MOT) remains challenging due to object occlusion, complex motions, and detection unreliability in crowded scenarios. We propose an enhanced MOT framework integrating and optimizing state-of-the-art components, specifically Improved Detection Confidence Boost (IDCBoost) and Track-Perspective-Based...
Trung Nghia Huynh, Chi Nhan Huynh, Jia-Ching Wang et al.· International Conference on...· 0 citations
This work proposes a cascade optimization framework that systematically enhances feature representation and refines multimodal fusion, and introduces the Multi-Scale Contextual Fusion Module (MSCF) to reduce alignment bias.