Skip to content
Preprint

MM-BEV: Enhancing Timeliness by Computing Where and When it Matters

Aug 2026 · 0 citations · 38 references
Computer Science Engineering

TL;DR

MM-BEV divides perception into mandatory work for safety-critical objects within braking distance of the ego vehicle and with short time-to-collision (TTC), and optional work for less urgent regions, which prioritizes mandatory work and reduces or sheds optional work under tight compute budgets.

Abstract

Multimodal bird's-eye-view (BEV) perception combines LiDAR depth accuracy with dense camera semantics, but its high computational cost and imperfect sensing conditions make real-time deployment challenging. Existing methods largely compress individual detectors and overlook three opportunities: structured sparsity within camera and LiDAR inputs, timing misalignment between modalities, and the fact that many detected objects do not affect the planner's immediate action. We present MM-BEV, a real-time multimodal BEV system guided by a simple principle: compute where and when it matters. MM-BEV divides perception into mandatory work for safety-critical objects within braking distance of the ego vehicle and with short time-to-collision (TTC), and optional work for less urgent regions. It prioritizes mandatory work and reduces or sheds optional work under tight compute budgets. MM-BEV integrates four mechanisms: (1) a criticality-ranked temporal ROI selector based on motion-extrapolated detections from prior frames; (2) sparse, ROI-aware feature extraction using shared-shape camera crops at context-adaptive resolution and ROI-aware LiDAR voxelization; (3) a latency-aware coordinator that adapts LiDAR sweeps, image resolution, and keyframes according to scene dynamics and TTC; and (4) an asynchronous scheduler that decouples sensing from inference and skips stale frames. On nuScenes, MM-BEV reduces inference latency by 1.96x and end-to-end latency by 2.93x, with no loss in geometry-critical recall and only a 0.2 percentage-point drop in safety-critical recall. On a Clearpath Husky A300 equipped with an Ouster-128 LiDAR, BEV cameras, and a Jetson AGX Orin, MM-BEV further reduces mean latency by 2.11x, demonstrating its potential for real-world autonomous systems.

View source

Similar papers

2026

GCAFormer-VO: A Geometry and Correspondence-Aware Transformer for Visual Odometry

Visual odometry (VO) estimates camera motion from image sequences and is essential for robotics, autonomous driving, and AR/VR. Robust VO remains challenging because large viewpoint changes and strong parallax make reliable cross-frame motion cues difficult to capture, especially in the presence of visual disturbances...

Jun-Qi Bao, Qing-Ying Wu, Jun Huang et al. · 0 citations
Open access Sep 2026

Ego-Motion-Aware Temporal Fusion in BEV Space for Multi-Modal 3D Object Detection

Multi-modal 3D object detection is critical for autonomous driving perception. While Bird’s Eye View (BEV) fusion methods effectively integrate LiDAR and camera features, they primarily focus on single-frame fusion and neglect temporal context. We propose CamT-BEV, a camera-temporal-enhanced BEV fusion framework for im...

Na Zhang, Edmundo Guerra, A. Grau · 0 citations
Nov 2026

MonoTemp3D: Lightweight Short-Horizon Temporal Enhancement for Streaming Monocular 3D Car Detection

Monocular 3D object detection (M3OD) is attractive for autonomous driving and resource-constrained robotic platforms due to its low hardware cost. While temporal cues can alleviate depth ambiguity, dense feature matching and explicit motion estimation introduce substantial latency and limit low-latency deployment. To a...

Jia-Ying Li, Tong Liu, Xiao-Dong Guo et al. · 0 citations
Preprint Sep 2026

Understanding Dynamic Scenes at Gigapixel Scale: Wide-Area Spatio-Temporal Perception from UAVs

UAV-borne imaging has advanced from megapixel to gigapixel sensors, shifting aerial perception from recognizing individual targets to understanding entire dynamic scenes. We characterize this demand as Wide-area Spatio-temporal Scene Understanding (WSTU), which requires wide-area coverage, per-target resolution, and te...

Yu-Hang Zhu, Mei-Yi Zhu, Yun-Kai Dang et al. · 0 citations
Preprint Sep 2026

DensePed-Lite: Quality-Aware Adaptive Detection for Dense Pedestrians under Occlusion

Pedestrian detection plays a crucial role in computer vision with applications in autonomous driving, surveillance, and public safety. However, real-world dense scenes bring severe challenges, including heavy occlusion, drastic scale variations, and strict real-time requirements. Existing lightweight detectors struggle...

Zian Wang, Ming-Zhe Liu, Chao-Yi Guo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.