Skip to content

DeGuNet: Depth-Guided Ultra-Compact Backbones for Efficient LiDAR-Camera 3D Detection

Jul 2026 · arXiv.org · Vol abs/2607.12419 · 0 citations · 59 references
Computer Science

TL;DR

DeGuNet is presented, an ultra-compact and plug-and-play image backbone explicitly designed for depth-guided representation learning that effectively aligns multi-view images with unstructured LiDAR depth while strictly preventing invalid-region contamination.

Abstract

In autonomous driving perception, the fusion of LiDAR and camera modalities has become the dominant paradigm for 3D object detection. However, current multi-modal frameworks heavily rely on massive visual backbones pretrained on 2D semantic tasks. This reliance introduces substantial parameter redundancy and a structural misalignment, as 2D priors are ill-equipped to handle the extreme sparsity of LiDAR projections required for Bird's-Eye-View geometry. To address this, we present DeGuNet, an ultra-compact and plug-and-play image backbone explicitly designed for depth-guided representation learning. By incorporating sparsity-aware feature extraction mechanisms, DeGuNet effectively aligns multi-view images with unstructured LiDAR depth while strictly preventing invalid-region contamination. Extensive experiments on the nuScenes dataset demonstrate DeGuNet's broad plug-and-play applicability and superior efficiency. When integrated into established baselines, it fundamentally eliminates architectural redundancy, reducing GPU memory consumption by up to 66.5% and achieving a 1.16x inference speedup. Concurrently, DeGuNet delivers up to a 6.20 absolute mAP gain, establishing a new paradigm for parameter-efficient multi-modal 3D perception.

View source

Similar papers

Preprint Sep 2026

DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models

We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation models (VFMs) to match monocular camera queries against a LiDAR map without modality-specific encoders. This enables robots and autonomous vehicles to robustly localize using only cameras within pre-built LiDAR map...

Yu-Hang Han, Youngseok Jang, Seungwon Roh et al. · 0 citations
Preprint Aug 2026

Cyclops: LiDAR as a Camera That Dreams in Color

Cyclops is proposed, a framework that translates sparse Non-Repetitive Scanning LiDAR intensity into RGB video, enabling camera-free inference for all-day perception tasks and mitigating inter-frame flickering.

Wei Gao, Jian Shu, Ming-Le Zhao et al. · 0 citations
Preprint Aug 2026

Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs

Map-Det3D is an online multi-view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB, suggesting that training reconstruction priors for detection is a practical route to stable metric 3D detection from monocular video.

Yung-Hsu Yang, Luigi Piccinelli, S. R. Bulò et al. · 1 citation
Preprint Sep 2026

Towards robust multimodal 3D object detection via visual foundation models

This work introduces SAM-AD, a domain-specific pretraining strategy that fine-tunes SAM on autonomous-driving imagery to extract feature representations with rich semantic information, and develops the Depth-Guided Wavelet Attention (DGWA) module, which suppresses high-frequency sensor noise while preserving critical c...

Zi-Ying Song, Lin Liu, Hong-Yu Pan et al. · 0 citations
Preprint Aug 2026

Geometry-Grounded Unified 3D Perception for Autonomous Driving

A Geometry-grounded Unified 3D Perception (GeoUP) framework that adapts the reconstruction-oriented latent of VGGT to calibrated, streaming multi-camera driving scenes and achieves SOTA performance across detection, occupancy, and depth estimation is presented.

Longfei Xu, Xiao-Hui Wang, Ze-Hao Huang et al. · 2 citations
Preprint Sep 2026

VDGS: Visibility-Driven Large-Scale 3D Gaussian Splatting for Aerial Scene Reconstruction

VDGS introduces visibility-driven statistics for scene anchors to quantify supervision strength and is leveraged for scene partitioning and for gradient compensation in under-optimized regions, thereby promoting balanced optimization across different regions.

Hao-Lin Yu, Jia-Dong Tang, Yi-Xian Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.