Nov 2026· IEEE Robotics and Automation Letters· Vol 11, pp. 12400-12407· 0 citations· 35 references
Abstract
Monocular 3D object detection (M3OD) is attractive for autonomous driving and resource-constrained robotic platforms due to its low hardware cost. While temporal cues can alleviate depth ambiguity, dense feature matching and explicit motion estimation introduce substantial latency and limit low-latency deployment. To alleviate this burden, we introduce MonoTemp3D, a lightweight dual-level temporal enhancement framework tailored to causal streaming query-based monocular 3D car detection. It efficiently exploits short-term temporal continuity and compact historical representations while preserving low-latency inference. To aggregate informative historical features without heavy pixel-wise matching, an asymmetric cross-scale temporal fusion module is proposed to employ sparse deformable attention as an implicit, task-driven temporal sampler. In addition, we introduce a context-aware query updating mechanism to propagate high-confidence historical object queries as instance-level priors and update them using current observations to stabilize 3D detection. Experiments on KITTI Tracking, KITTI 3D and nuScenes show that MonoTemp3D outperforms representative image-based and video-based M3OD methods. MonoTemp3D achieves inference latencies of 34 ms on an RTX 3090 GPU and 43.5 ms on a Jetson Orin NX (TensorRT, FP16).
Monocular temporal 3D detection aims to detect objects in 3D, given a monocular video. Query-based 3D detectors unify detection and cross-view association, but their learnable queries fit the spatial distribution of the training data (e.g., field-of-view). We show that this issue is especially severe when these models...
M. Kotb, Johannes Meier, Christoph Reich et al.· 0 citations
LiFR v2 is presented, a unified propagation-completion-memory framework for causal anytime and streaming dense prediction from an RGB keyframe and events and introduces an Event-Guided Completion Module (EGCM) and a History Retrieval Module (HRM) to reuse completed representations across successive queries.
Tao Wan, Xiao-Shan Wu, Yi-Fei Yu et al.· 0 citations
Self-supervised monocular depth estimation (MDE) eliminates the reliance on expensive ground-truth depth annotations and has emerged as a powerful approach for a wide range of vision applications. However, current lightweight networks are hampered by two critical challenges: the limited representation capacity of stati...
Dong-Liang Wang, Ming Jin, Xiang-Qian Fang et al.· International Conference on...· 0 citations
Transformer-based object trackers are renowned for their strong performance, yet dense token processing often leads to prohibitive computational cost, limiting real-time deployment on edge devices. While recent works explore token pruning to reduce computation, they often stop short of an end-to-end sparse pipeline, as...
Qingmao Wei, Fa-Gui Liu, Dengke Zhang et al.· 0 citations
We propose a scan-by-scan framework for LiDAR 3D object detection that eliminates redundant re-encoding of overlapping historical scans by propagating temporal context through internal memory. This formulation raises two fundamental challenges: inter-step feature misalignment and long-horizon degradation that conventio...
Eisho Tsuji, Ryuhei Hamaguchi, Masaki Onishi et al.· IEEE Robotics and Automation...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.