Sep 2026· AI Photonics Technology Symposium· Vol 14312, pp. 1431204 - 1431204-6· 0 citations· 14 references
Engineering
TL;DR
RGD-Net, a new depth estimation network that incorporates RGB geometric priors that can capture long-range spatial dependencies and reduce local noise interference, is suggested, and experimental results indicate that RGD-Net is promising for range-gated 3D imaging in autonomous perception tasks.
Abstract
Existing deep learning-based 3D range-gated imaging methods, like Gated2Depth, still lack the ability to fully exploit color cues, have high sensitivity to specular noise, and produces undesirable local artifacts in autonomous driving situations. To overcome these limitations, we suggest RGD-Net, a new depth estimation network that incorporates RGB geometric priors. The network employs three consecutive range-gated frames as its input. Synchronized RGB images are projected and aligned with the gated coordinate system using camera calibration parameters to provide geometric priors such as scene edges, texture gradients, and spatial structures. Based on the encoder-Transformer-decoder architecture, this framework employs a cross-modal fusion strategy with a cross-attention mechanism to effectively combine depth features from gated images with edge and texture information from RGB images. Using the global modeling ability of Transformer, our model can capture long-range spatial dependencies and reduce local noise interference. Edge alignment and relative depth consistency constraints are further designed to facilitate the learning of depth boundaries and scene geometric structures. Experimental results on the Gated2Depth dataset demonstrate that compared with the baseline, the proposed method reduces MAE by 42.84% and RMSE by 29.00% in daytime scenarios, while boosting δ1 accuracy by 18.89 percentage points. The generated dense depth maps exhibit clearer boundaries, effectively correcting erroneous predictions in depth-discontinuous regions and improving the integrity and accuracy of object contours. Experimental results indicate that RGD-Net is promising for range-gated 3D imaging in autonomous perception tasks.
Results confirm the effectiveness of combining transformer-based global encoding with lightweight convolutional decoding for high-quality, real-time monocular depth estimation.
Multimodal perception integrating light detection and ranging (LiDAR) and cameras has become a key paradigm for 3-D object detection, as it leverages both geometric structure and semantic information. However, in real-world autonomous driving scenarios, calibration errors, adverse weather, and sensor degradation can in...
Hui-Lin Huang, Yan Bai, Peng-Yuan Wang et al.· IEEE Sensors Journal· 0 citations
A dual representation-based LFVS method that employs deformable convolutional and Deep Residual Channel Attention (DRCA) networks that achieves state-of-the-art performance on synthetic and real-world LF benchmarks.
Muhammad Zubair, Paulo J. L. Nunes, Caroline Conti et al.· IEEE Open Journal of Signal...· 0 citations
It is concluded that future industrial deployment on edge-computing platforms will rely on a synergy between lightweight network architectures and multi-sensor fusion and self-supervised frameworks.
A hybrid self-supervised model incorporating a Skip Attention Mechanism to minimize semantic inconsistencies between encoder and decoder, and a mixed pooling method combining max and average pooling to address over-smoothing and contextual information loss is proposed.
Afrasiab Khan, Tahir Nawaz, M. Asaduzzaman et al.· Signal, Image and Video Proc...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.