Aug 2026· Neural Networks· Vol 205 Pt B, pp.
109451
· 0 citations· 60 references
Medicine
TL;DR
A Graph-driven Contextual Synergy Network (GCS3D), which is designed to systematically enhance point representations across both semantic and geometric dimensions, and introduces a Graph-guided Geometric Consistency Interaction module for contextual correlation modeling.
Abstract
3D object detection stands as a pivotal task in scene understanding. However, two primary bottlenecks constrain current methodologies: semantic ambiguity arising from spatial misalignment during cross-modal fusion, and the inadequate contextual representation of individual candidate points within complex scenes. To address these challenges, this paper presents a Graph-driven Contextual Synergy Network (GCS3D), which is designed to systematically enhance point representations across both semantic and geometric dimensions. Specifically, the proposed method incorporates a Semantic Representation Rectification (G-SRR) module for cross-modal representation enhancement. By performing region-level semantic aggregation based on 3D neighborhoods to mitigate projection bias, this module achieves robust cross-modal fusion through a Spatial-aware Gating Mechanism (SGM) that adaptively regulates visual feature injection. Regarding contextual correlation modeling, the framework introduces a Graph-guided Geometric Consistency Interaction (G-GCI) module. By constructing a local topology graph among anchors and executing position-aware feature interaction, this module facilitates the aggregation of complementary neighborhood information, thereby bolstering the feature consistency and discriminability of anchor representations. Furthermore, a Spatial-Scale Aware Assigner (SSA-Assigner) is utilized to dynamically allocate supervision signals based on prediction quality, fully exploiting the performance potential inherent in the enhanced anchor representations. Extensive experiments on the SUN RGB-D and ScanNet V2 datasets demonstrate that GCS3D achieves superior results with mAP@0.25 scores of 70.39 and 73.86 respectively, validating the effectiveness and robustness of the proposed strategy in complex indoor scenes.
This work proposes AMCA, an unbiased SGG framework integrating adaptive multi-prototype learning with cross-modal alignment, which achieves consistently competitive performance across multiple SGG tasks, with particularly strong improvements on the unbiased mR@K metric.
Jinhao Fan, Yuanhao Xi, Chuanping Hu et al.· Journal of King Saud Univers...· 0 citations
A 3D Local- Global Linear Attention Mechanism (LG-LAM) is devised that efficiently captures long-range contextual information with linear complexity, enabling a comprehensive understanding of the 3D scene without heavy computational burdens.
Jie Li, Jia-Heng Xu, Laiyan Ding et al.· International Conference on...· 0 citations
This work proposes a segmentation framework guided by Contrastive Language–Image Pre-training (CLIP) that enriches sparse 3D tokens with vision–language semantic priors and introduces a decoupled CLIP-induced semantic residual that forms semantic-geometric attention biases for local window attention.
Prior-SG achieves state-of-the-art semantic region segmentation accuracy compared to recent baselines, robustly delineates distant functional boundaries in the absence of physical walls, and uniquely provides zero-shot ontological flexibility, enabling the robot to entirely restructure its spatial partitioning based on...
G. Tonetti, Laurent Kneip, Abel Gawel et al.· 0 citations
A novel Coarse Map-guided Multi-scale Interaction Network (CMNet), which employs a dual-branch structure to enhance both semantic context and fine-grained details, and introduces a progressive fusion mechanism to generate coarse priors that guide subsequent feature refinement.
Meng-Ju Lu, Mingyong Pang· Multimedia tools and applica...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.