Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· pp. 926-934· 0 citations· 39 references
TL;DR
This work designs a semantic structure-based feature refinement module to adaptively modulate subject-centric and contextual visual features and incorporates a visual semantic consistency verification module, which leverages a subject-aware contrastive learning strategy to constrain and verify the semantic correspondence between the visual prediction and subject-level textual representation.
Abstract
Visual grounding aims to localize target objects based on natural language descriptions, and the core challenge lies in the cross-modal gap, which is partly caused by the significant differences in semantic structure between language and vision. Existing methods typically rely on holistic sentence-level semantic representations to modulate visual features, while overlooking the inherent structure of textual prompts. In this work, we propose a Syntactic Structure-guided Visual Grounding framework, referred to as SSVG. Specifically, to inject syntactic priors into unstructured visual representations, we design a semantic structure-based feature refinement module to adaptively modulate subject-centric and contextual visual features. To perform cross-modal alignment, we further incorporate a visual semantic consistency verification module, which leverages a subject-aware contrastive learning strategy to constrain and verify the semantic correspondence between the visual prediction and subject-level textual representation, thereby enhancing the model's robustness against semantically similar distractors. Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art methods, and detailed analyses verify the effectiveness of each component.
Introduction Word grounding refers to the ability of agents to associate linguistic terms with the perceptual concepts they represent, such as color, shape, and spatial features. This study quantitatively evaluates how syntactic and semantic information affects word-grounding performance. Methods We evaluated five conf...
Saima Shaukat, A. Aly, Anouar Chibani et al.· Frontiers in Artificial Inte...· 0 citations
Image description generation faces the challenge of insufficient visual semantic structuring in cross-linguistic scenarios. Inspired by systemic functional linguistics and cognitive grammar, this paper designs a Perceptual Visual Context Encoder (PVCE), which transforms pixel signals into semantic propositions with ima...
Hui Li, Hui Chen· International Workshop on Ar...· 0 citations
ReVA is proposed, a region-aware VQA model that employs a frozen CLIP ViT-L/14 Vision Transformer (ViT) and a Qwen2.5-7B-Instruct large language model (LLM) connected through a dual bridge that aligns both whole-image and region-level representations with the LLM's embedding space.
Think with Structured Grounding (TwSG), a novel fine-grained image perception framework designed to internalize complex images's tool-use capabilities within the model, is proposed, endowing models with native fine-grained region description and flexible reasoning capabilities.
Chang-Jiang Jiang, Qian-Nian Zhao, Lei Xin et al.· 2 citations
Generalized Visual Grounding (GVG) task aims to localize targets in an image based on referring expressions, extends the classical visual grounding paradigm by integrating multi-target and non-target scenarios. Previous methods typically rely on global semantic matching or coarse-grained region interactions for localiz...
FSG-AID is presented, which integrates fine-grained semantic guidance with an attribute-aware iterative decoder and jointly exploits visual and language features to mine attribute semantics, initialize the target query, and iteratively refine the target representation.
Xiya Bu, Yu Liu, Jizhe Yu et al.· Multimedia Systems· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.