Skip to content

3D Dynamic Object Detection with Multi-Modal Temporal Fusion

Sep 2026 · SAE technical paper series · 0 citations · 1 references

Abstract

This research aims to address the critical challenge of accurately detecting and estimating the state of dynamic objects in autonomous driving. Traditional 3D object detection methods often struggle with motion perception, particularly in velocity estimation, due to the lack of information in single frame perception. We propose a novel framework that enhances the BEV representation with temporal modeling. The core of our method is a two-stage temporal fusion process. First, we align historical BEV features to the current coordinate frame to eliminate the interference of ego-motion. Subsequently, a dedicated temporal fusion encoder, architected with residual connections and a Feature Pyramid Network, refines the aligned multi-frame BEV features to capture complex motion patterns and improve multi-scale object representation. This approach directly tackles the problem of motion decoupling. By aligning features, we disentangle object motion from ego-motion. The temporal fusion encoder then mitigates the positional ambiguity of moving objects in the fused BEV space, a common issue in simple feature concatenation, leading to more robust detection. We built a dataset following the structure of the nuScenes dataset, using data collected from an autonomous driving simulation platform. The evaluation results on our simulation dataset demonstrate that the proposed temporal module achieves a 13.0% improvement in NDS score and a substantial 29.7% reduction in velocity error (mAVE). These results demonstrate that our temporal fusion strategy effectively enhances 3D detection accuracy in dynamic scenarios.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.