Skip to content
Conference

Asymmetric Cross-Modal Modulation for Spatial Relation Recognition

Aug 2026 · 2026 12th International Conference on Big Data and Information Analytics (BigDIA) · pp. 515-522 · 0 citations · 32 references

Abstract

Recognizing spatial geometric relationships between objects is crucial for machines to understand the physical world. Although introducing the depth modality can effectively mitigate the ambiguities caused by 2D visual projection, existing cross-modal fusion paradigms heavily rely on computationally expensive late-stage cross-attention decoders or blind symmetric feature interactions, ignoring the fundamental physical heterogeneities between RGB and depth modalities. In this paper, we propose Entity-Aware Asymmetric Mutual Modulation Vision Transformer (EAM-ViT), a streamlined and physical-intuition-driven architecture. Initially, we employ a shared-weight isomorphic encoder to process dual-stream inputs, minimizing parameter overhead while enforcing strict early alignment of multi-modal features. The core of our framework is the EAM module, which enables the two modalities to perform directional, asymmetric guidance based on their unique physical characteristics. Specifically, in the spatial pathway, depth features leverage Spatial Prior Modulation (SPM) to force the visual network to precisely focus on physical interaction boundaries and filter background noise; in the channel pathway, pure visual semantics utilize Semantic-Driven Channel Calibration (SCC) to adaptively recalibrate 3D geometric depth channels. Extensive experiments on the SpatialSense and SpatialSense+ datasets demonstrate that our model significantly outperforms state-of-the-art methods.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.