A novel framework that mines target knowledge using the strong reasoning capacity of Multimodal Large Language Models (MLLMs) and employs it to enhance FSS, which shows promising results and largely surpasses existing methods.
Abstract
Few-shot segmentation (FSS) aims to segment unseen object categories with a few (e.g., one or five) labeled examples, enabling efficient adaptation to novel classes. Conventional models typically rely on appearance-based visual matching between support and query images for segmentation. While straightforward, these methods often struggle to handle significant appearance discrepancies and occlusions in the query image due to insufficient target knowledge. To mitigate this, we introduce a novel framework that mines target knowledge using the strong reasoning capacity of Multimodal Large Language Models (MLLMs) and employs it to enhance FSS. Specifically, building on SAM 2, our method, named MK-FSS, exploits two forms of complementary knowledge derived from a query image by an MLLM for FSS, including spatial knowledge, which provides a spatial prior indicating the potential target location, and semantic knowledge, which describes the target using text. The spatial knowledge is first encoded into a memory representation, and then resulting memory is integrated with the support-guided memory feature from query image through a carefully designed dual-memory debate-fusion (DMDF) module, yielding a more robust target memory feature. In parallel, the semantic knowledge is encoded into the textual feature, which is fused with multi-scale query features via a progressive cross-modal prompt generator (PCPG), producing a target-aware multimodal prompt for segmentation. Working together, the dual-memory feature and the multimodal prompt provide a comprehensive representation of the target, enabling more robust segmentation. In our extensive experiments, MK-FSS shows promising results and largely surpasses existing methods. Code will be released.
A multimodal Graph-based Fusion (MGF) framework that learns visually grounded semantic representations to enhance FSOR performance and achieves superior open-set recognition and competitive closed-set classification performance.
This work formulates active learning for VG under the realistic setting where only raw images are available without accompanying text, and introduces Referred Region Ambiguity, a new acquisition function that measures whether the model's confidence collapses onto a single region or disperses across multiple candidates.
Junbeom Hong, Seonghoon Yu, Hyungsik Jung et al.· 0 citations
A text-guided Zero-Shot Localization framework for unseen object categories (ZSOL) for addressing the aforementioned challenges, which can be guided by prompt words to identify and localize unseen object categories in images by transferring localization knowledge learned from supervised base categories.
Few-shot segmentation (FSS) remains a significant challenge due to the scarcity of annotated data and the need for precise object localization across novel classes. Existing approaches often rely on single-backbone architectures and coarse priors, which struggle to capture detailed semantics and precise spatial alignme...
Guo-Hua Geng, Xiao-Feng Wang· International Conference on...· 0 citations
Few-shot classification (FSC) aims to emulate the human ability to rapidly learn new concepts from a handful of examples. Large Vision-Language Models (LVLMs), with their rich prior knowledge and powerful visio-linguistic understanding capabilities, are emerging as a highly promising paradigm for FSC. This paper invest...
Chuan-Yi Zhang, Fan Liu, Yi Xu et al.· IEEE Transactions on Image P...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.