Skip to content
Conference Open access

Syntactic Structure-Guided Visual Grounding with Subject-Centric Feature Enhancement and Verification

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · pp. 926-934 · 0 citations · 39 references

TL;DR

This work designs a semantic structure-based feature refinement module to adaptively modulate subject-centric and contextual visual features and incorporates a visual semantic consistency verification module, which leverages a subject-aware contrastive learning strategy to constrain and verify the semantic correspondence between the visual prediction and subject-level textual representation.

Abstract

Visual grounding aims to localize target objects based on natural language descriptions, and the core challenge lies in the cross-modal gap, which is partly caused by the significant differences in semantic structure between language and vision. Existing methods typically rely on holistic sentence-level semantic representations to modulate visual features, while overlooking the inherent structure of textual prompts. In this work, we propose a Syntactic Structure-guided Visual Grounding framework, referred to as SSVG. Specifically, to inject syntactic priors into unstructured visual representations, we design a semantic structure-based feature refinement module to adaptively modulate subject-centric and contextual visual features. To perform cross-modal alignment, we further incorporate a visual semantic consistency verification module, which leverages a subject-aware contrastive learning strategy to constrain and verify the semantic correspondence between the visual prediction and subject-level textual representation, thereby enhancing the model's robustness against semantically similar distractors. Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art methods, and detailed analyses verify the effectiveness of each component.

Read PDF

Similar papers

Open access Sep 2026

The effect of syntactic and semantic information on word grounding through visual perception

Introduction Word grounding refers to the ability of agents to associate linguistic terms with the perceptual concepts they represent, such as color, shape, and spatial features. This study quantitatively evaluates how syntactic and semantic information affects word-grounding performance. Methods We evaluated five conf...

Saima Shaukat, A. Aly, Anouar Chibani et al. · 0 citations
Conference Aug 2026

Multimodal Image Description Generation and Alignment Mechanism Integrating Visual Context

Image description generation faces the challenge of insufficient visual semantic structuring in cross-linguistic scenarios. Inspired by systemic functional linguistics and cognitive grammar, this paper designs a Perceptual Visual Context Encoder (PVCE), which transforms pixel signals into semantic propositions with ima...

Hui Li, Hui Chen · 0 citations
#computer vision Preprint Aug 2026

ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering

ReVA is proposed, a region-aware VQA model that employs a frozen CLIP ViT-L/14 Vision Transformer (ViT) and a Qwen2.5-7B-Instruct large language model (LLM) connected through a dual bridge that aligns both whole-image and region-level representations with the LLM's embedding space.

A. Senthil · 0 citations
Preprint Aug 2026

Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding

Think with Structured Grounding (TwSG), a novel fine-grained image perception framework designed to internalize complex images's tool-use capabilities within the model, is proposed, endowing models with native fine-grained region description and flexible reasoning capabilities.

Chang-Jiang Jiang, Qian-Nian Zhao, Lei Xin et al. · 2 citations
Preprint Aug 2026

Semantic-Spatial Discriminability Enhancement for Generalized Visual Grounding

Generalized Visual Grounding (GVG) task aims to localize targets in an image based on referring expressions, extends the classical visual grounding paradigm by integrating multi-target and non-target scenarios. Previous methods typically rely on global semantic matching or coarse-grained region interactions for localiz...

Kai Lei, Xu-Yao Zhang · 0 citations
Sep 2026

Enhancing target identification and query discrimination for visual grounding

FSG-AID is presented, which integrates fine-grained semantic guidance with an attribute-aware iterative decoder and jointly exploits visual and language features to mine attribute semantics, initialize the target query, and iteratively refine the target representation.

Xiya Bu, Yu Liu, Jizhe Yu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.