Skip to content

Enhancing Multimodal Compositional Understanding of Vision-Language Models With Semantic Decoupling and Feature Coupling

2026 · IEEE Signal Processing Letters · Vol 33, pp. 3102-3106 · 0 citations · 39 references

Abstract

Vision-Language Models (VLMs) have achieved remarkable success across various downstream tasks. However, their compositional understanding of complex attributes and relations remains a significant challenge. While existing methods leverage compositional datasets and contrastive learning, they are often compromised by various interferences such as textual ambiguity and synthetic artifacts. To address these issues, this paper proposes DC-CLIP, a framework designed to extract and leverage fine-grained semantic primitives from image-text pairs. Specifically, a Semantic Decoupling Module (SDM) is introduced to decompose global inputs into fine-grained entities and their corresponding local regions. Then, these primitives are fed into the encoders alongside the original data, with feature integration facilitated by a Feature Coupling Module (FCM). Furthermore, a Decoupling Contrastive Loss (DCL) is proposed to strengthen the representation learning of critical semantic features. Extensive experiments on ARO and SugarCrepe demonstrate that DC-CLIP significantly outperforms state-of-the-art methods in compositional understanding tasks.

View source