Rethinking Temporal Feature Propagation for Efficient 3D Object Detection
Abstract
We propose a scan-by-scan framework for LiDAR 3D object detection that eliminates redundant re-encoding of overlapping historical scans by propagating temporal context through internal memory. This formulation raises two fundamental challenges: inter-step feature misalignment and long-horizon degradation that conventional short-window evaluation can obscure. To address the former, we introduce Cross-Attention Motion Transformation (X-MoT), a motion-aware recurrent update module that aligns propagated memory with the current observation through coarse egomotion compensation followed by cross-attention-based refinement. To address the latter, we introduce sequential evaluation under continuous inference without state resets, and show that long-sequence training and Union-QK further improve robustness in this setting. Under conventional evaluation, X-MoT achieves a favorable accuracy-latency trade-off against multi-scan baselines. At the same time, sequential evaluation reveals substantial long-horizon degradation in memory-based detectors that is obscured by conventional evaluation. Taken together, these results show that scan-by-scan detection can deliver competitive accuracy at low latency, provided that its long-horizon behavior is explicitly evaluated and stabilized.