Skip to content

Graph-driven contextual synergy network for robust 3D object detection.

Aug 2026 · Neural Networks · Vol 205 Pt B, pp. 109451 · 0 citations · 60 references
Medicine

TL;DR

A Graph-driven Contextual Synergy Network (GCS3D), which is designed to systematically enhance point representations across both semantic and geometric dimensions, and introduces a Graph-guided Geometric Consistency Interaction module for contextual correlation modeling.

Abstract

3D object detection stands as a pivotal task in scene understanding. However, two primary bottlenecks constrain current methodologies: semantic ambiguity arising from spatial misalignment during cross-modal fusion, and the inadequate contextual representation of individual candidate points within complex scenes. To address these challenges, this paper presents a Graph-driven Contextual Synergy Network (GCS3D), which is designed to systematically enhance point representations across both semantic and geometric dimensions. Specifically, the proposed method incorporates a Semantic Representation Rectification (G-SRR) module for cross-modal representation enhancement. By performing region-level semantic aggregation based on 3D neighborhoods to mitigate projection bias, this module achieves robust cross-modal fusion through a Spatial-aware Gating Mechanism (SGM) that adaptively regulates visual feature injection. Regarding contextual correlation modeling, the framework introduces a Graph-guided Geometric Consistency Interaction (G-GCI) module. By constructing a local topology graph among anchors and executing position-aware feature interaction, this module facilitates the aggregation of complementary neighborhood information, thereby bolstering the feature consistency and discriminability of anchor representations. Furthermore, a Spatial-Scale Aware Assigner (SSA-Assigner) is utilized to dynamically allocate supervision signals based on prediction quality, fully exploiting the performance potential inherent in the enhanced anchor representations. Extensive experiments on the SUN RGB-D and ScanNet V2 datasets demonstrate that GCS3D achieves superior results with mAP@0.25 scores of 70.39 and 73.86 respectively, validating the effectiveness and robustness of the proposed strategy in complex indoor scenes.

View source

Similar papers

Open access Aug 2026

AMCA-SGG: Adaptive multi-prototype learning and cross-modal alignment for unbiased scene graph generation

This work proposes AMCA, an unbiased SGG framework integrating adaptive multi-prototype learning with cross-modal alignment, which achieves consistently competitive performance across multiple SGG tasks, with particularly strong improvements on the unbiased mR@K metric.

Jinhao Fan, Yuanhao Xi, Chuanping Hu et al. · 0 citations
Conference Aug 2026

Enhancing 3D semantic scene completion via efficient attention and feature augmentation

A 3D Local- Global Linear Attention Mechanism (LG-LAM) is devised that efficiently captures long-range contextual information with linear complexity, enabling a comprehensive understanding of the 3D scene without heavy computational burdens.

Jie Li, Jia-Heng Xu, Laiyan Ding et al. · 0 citations
Open access Sep 2026

Vision–language guided semantic-geometric transformer for memory-efficient 3D scene understanding

This work proposes a segmentation framework guided by Contrastive Language–Image Pre-training (CLIP) that enriches sparse 3D tokens with vision–language semantic priors and introduces a decoupled CLIP-induced semantic residual that forms semantic-geometric attention biases for local window attention.

Li-Cheng Liu, Yu Li, Fu-Yong Liu · 0 citations
Preprint Aug 2026

Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments

Prior-SG achieves state-of-the-art semantic region segmentation accuracy compared to recent baselines, robustly delineates distant functional boundaries in the absence of physical walls, and uniquely provides zero-shot ontological flexibility, enabling the robot to entirely restructure its spatial partitioning based on...

G. Tonetti, Laurent Kneip, Abel Gawel et al. · 0 citations
Aug 2026

CMNet: a coarse maps guided multi-scale interaction network for camouflaged object detection

A novel Coarse Map-guided Multi-scale Interaction Network (CMNet), which employs a dual-branch structure to enhance both semantic context and fine-grained details, and introduces a progressive fusion mechanism to generate coarse priors that guide subsequent feature refinement.

Meng-Ju Lu, Mingyong Pang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.