Skip to content
Open access

EAEPnet: multi-level representation enhancement and weighted box fusion network for multi-modal 3D object detection

Aug 2026 · Measurement science and technology · Vol 37 · 0 citations · 39 references
Physics

TL;DR

A aggregated Euclidean distance weighted box fusion method, which aggregates complementary information from multiple candidate boxes during post-processing to improve bounding-box selection and localization accuracy, and a hybrid deformable half-conv (HDHC) module that jointly enhances global and local feature representations through hierarchical offset prediction and local neighborhood attention are proposed.

Abstract

With the rapid advancement of intelligent driving, LiDAR and cameras provide complementary geometric and semantic information. Therefore, fusing these two sensors is a mainstream approach for 3D object detection. However, existing methods still face two key challenges: effectively exploiting complementary information among candidate boxes during post-processing and balancing detection accuracy with computational efficiency. To address these challenges, we first propose an aggregated Euclidean distance weighted box fusion (AED-WBF) method, which aggregates complementary information from multiple candidate boxes during post-processing to improve bounding-box selection and localization accuracy. We further develop a hybrid deformable half-conv (HDHC) module that jointly enhances global and local feature representations through hierarchical offset prediction and local neighborhood attention. By integrating half-conv with a separable self-attention mechanism, HDHC reduces computational complexity while maintaining detection accuracy. Based on AED-WBF and HDHC, we construct EAEPNet, an efficient multilevel LiDAR–camera fusion network for 3D object detection. Extensive experiments are conducted on the KITTI and nuScenes datasets. On the KITTI test set, EAEPNet improves the mean average precision (mAP) by 2.73% over the baseline network. On nuScenes, EAEPNet achieves a mAP of 72.5% and an nuScenes detection score (NDS) of 74.4% on the validation set, as well as a mAP of 73.2% and an NDS of 75.3% on the test set. These results validate the effectiveness of EAEPNet in multi-sensor 3D object detection and spatial measurement. Its strong performance across multiple datasets further demonstrates its potential for intelligent driving and real-time high-precision spatial measurement. The code is available at: https://github.com/juanmao73/EAEPNet.

Read PDF

Similar papers

Aug 2026

Fadet: a fusion-aware 3D detection network with cascaded feature enhancement for small object detection in autonomous driving

This work proposes a cascade optimization framework that systematically enhances feature representation and refines multimodal fusion, and introduces the Multi-Scale Contextual Fusion Module (MSCF) to reduce alignment bias.

Chang-Hong Yu, Shaoshi Luo, Wen-Li Shen · 0 citations
Sep 2026

Class-Enhanced Multi-Sampling and Multi-Level Graph Attention for 3-D Object Detection

Three-dimensional object detection is a key component of the perception module in autonomous driving systems. Compared to camera images, LiDAR point clouds provide richer spatial information, such as detailed structural and geometric cues of objects. However, existing 3D object detection methods face two major challeng...

Xiangyang Wu, Ji-Tao Pan, Qing-Long Jiao et al. · 0 citations
Conference Aug 2026

A multimodal BEV 3D object detection method with depth uncertainty and geometric saliency

High-precision perception is fundamental to safe autonomous driving, and BEV-based 3D object detection via lidar-camera fusion plays a crucial role in improving detection accuracy and robustness. To address insufficient feature representation, spatial misalignment, and the limitations of static fusion strategies, this...

Jie Hu, Xinghao Cheng, Shuaidi He et al. · 0 citations
Open access Sep 2026

Two-Stage 3D Object Detection Architecture via Integrated Attention Mechanism and Voxel Aggregation

In the realm of autonomous driving, 3D object detection based on LiDAR point clouds has emerged as a pivotal technology. To enhance the accuracy and efficiency of 3D object detection, this paper introduces Spatial-Channel Attention Guided with Gumbel Subset Sampling and Context Fusion RCNN (SCAGCF-RCNN), a two-stage...

Hong-Xu Li, Shu-Yi Zhou, Yi-Tao Lu et al. · 0 citations
Aug 2026

Ecf3dmot: enhanced centerpoint framework for 3D object detection and tracking with LiDAR

Results validate the effectiveness of the proposed novel 3D object detection and tracking framework, termed ECF3DMOT, in advancing 3D object detection and tracking for autonomous driving.

Xiaojuan Peng, Fei Teng, Tiankai Chen et al. · 0 citations
Open access Aug 2026

MVXCC-NET: Cross-modal 3D detection of occluded objects based on dual-path information complementation and regional weight modeling

MVXCC-NET is presented, a cross-modal 3D detection network for occluded objects based on dual-path information complementation and regional weight modeling, which improves the utilization efficiency of fused features, allowing visual semantic information and spatial geometric information to support each other.

Jin Qi, Jian Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.