Multimodal fusion of cameras and millimeter-wave radars is critical for robust all-weather object detection and ensuring vehicle safety in real-world autonomous driving vehicles. However, existing radar-camera fusion methods that rely on a unified Bird’s Eye View (BEV) representation often suffer from information loss and limited cross-modal interaction. To address these limitations, a query-based multimodal fusion framework, termed SRCDet, is proposed for camera-4D radar fusion. The framework processes features in parallel across both BEV and Perspective View spaces, where deformable attention is employed to achieve dynamic cross-view alignment. By integrating radar attribute features, a local–global dual-branch query generation mechanism is designed to produce high-quality 3D detection proposals. Furthermore, a graph neural network-based cross-fusion module is introduced to model complex inter-feature relationships through a heterogeneous interaction graph. Extensive experiments on the OmniHD-Scenes and NuScenes datasets demonstrate that SRCDet achieves consistent improvements across nearly all metrics and low error rates in clear and adverse weather conditions, highlighting its practical adaptability to automotive-grade systems and effectiveness in safety-critical real-world autonomous driving scenarios.
Wen-Jin Ai, Lianqing Zheng, Long Yang et al.· Measurement science and tech...· 0 citations
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches and spectral variations, alongside textual logic drift from unconstrained, variable user-input granularities. To mitigate these bottlenecks, this paper establishes the first cross-domain RRSIS benchmark, designated as the Vaihingen-Potsdam Referring (VPRef) dataset, comprising 46,972 language-image-annotation triplets organized into a three-tier linguistic hierarchy. Building upon this benchmark, we develop a tailored parameter-efficient domain adaptation baseline anchored on the Segment Anything Model (SAM3) via Low-Rank Adaptation (LoRA). Our framework counteracts visual distribution discrepancies through pseudo-label-driven self-training and addresses textual logic drift via random multi-granularity text prompt mixing. Crucially, the distribution of empirical metrics across ablative variants suggests a potential decoupling between cross-modal semantic robustification and visual domain alignment, demonstrating that linguistic variance drives fine-grained semantic invariance while pseudo-label propagation governs macro-scale spatial grid alignment. Extensive benchmarks demonstrate the proposed framework achieves superior cross-domain segmentation boundaries while modifying merely 1.08\% of the foundational parameter footprint, establishing a robust baseline for future multi-modal remote sensing domain adaptation research. The dataset and code will be available at https://github.com/quanweiliu/VPRef.
Quan-Wei Liu, Tao Huang, Jia-Qi Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.