Skip to content
Open access

Ego-Motion-Aware Temporal Fusion in BEV Space for Multi-Modal 3D Object Detection

Sep 2026 · Electronics · 0 citations · 8 references

Abstract

Multi-modal 3D object detection is critical for autonomous driving perception. While Bird’s Eye View (BEV) fusion methods effectively integrate LiDAR and camera features, they primarily focus on single-frame fusion and neglect temporal context. We propose CamT-BEV, a camera-temporal-enhanced BEV fusion framework for improved multi-modal 3D object detection. Our key insight is that temporal modeling is particularly critical for the camera branch to resolve monocular depth ambiguity and object occlusion, while single-frame LiDAR representation already provides accurate instantaneous geometry. We thus propose a camera-centric temporal enhancement module via ego-motion warping and ConvLSTM temporal encoding. Extensive experiments on the nuScenes dataset demonstrate that CamT-BEV achieves competitive perception performance, attaining 0.6971 NDS and 0.6683 mAP, with notable relative AP gains on challenging categories such as bicycles (+27.3%) and motorcycles (+7.66%) evaluated under category-level mAP (averaged across 0.5 m to 4.0 m distance thresholds). Furthermore, evaluations under fog and miss-beam conditions in nuScenes-C confirm its improved robustness against specific visual and sensor degradations. Crucially, these gains are achieved with low additional computational and memory overhead, demonstrating that targeted camera-temporal fusion is a practical solution for 3D perception.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.