Skip to content
Preprint

Exploiting Target Knowledge from MLLMs for Robust Few-Shot Segmentation

Sep 2026 · 0 citations · 55 references
Computer Science

TL;DR

A novel framework that mines target knowledge using the strong reasoning capacity of Multimodal Large Language Models (MLLMs) and employs it to enhance FSS, which shows promising results and largely surpasses existing methods.

Abstract

Few-shot segmentation (FSS) aims to segment unseen object categories with a few (e.g., one or five) labeled examples, enabling efficient adaptation to novel classes. Conventional models typically rely on appearance-based visual matching between support and query images for segmentation. While straightforward, these methods often struggle to handle significant appearance discrepancies and occlusions in the query image due to insufficient target knowledge. To mitigate this, we introduce a novel framework that mines target knowledge using the strong reasoning capacity of Multimodal Large Language Models (MLLMs) and employs it to enhance FSS. Specifically, building on SAM 2, our method, named MK-FSS, exploits two forms of complementary knowledge derived from a query image by an MLLM for FSS, including spatial knowledge, which provides a spatial prior indicating the potential target location, and semantic knowledge, which describes the target using text. The spatial knowledge is first encoded into a memory representation, and then resulting memory is integrated with the support-guided memory feature from query image through a carefully designed dual-memory debate-fusion (DMDF) module, yielding a more robust target memory feature. In parallel, the semantic knowledge is encoded into the textual feature, which is fused with multi-scale query features via a progressive cross-modal prompt generator (PCPG), producing a target-aware multimodal prompt for segmentation. Working together, the dual-memory feature and the multimodal prompt provide a comprehensive representation of the target, enabling more robust segmentation. In our extensive experiments, MK-FSS shows promising results and largely surpasses existing methods. Code will be released.

View source

Similar papers

Sep 2026

Multimodal graph-based fusion via image descriptions for few-shot open-set recognition.

A multimodal Graph-based Fusion (MGF) framework that learns visually grounded semantic representations to enhance FSOR performance and achieves superior open-set recognition and competitive closed-set classification performance.

Xilang Huang, Seon-Han Choi · 0 citations
#artificial intelligence Preprint Aug 2026

Cost-efficient Active Learning for Referring Image Segmentation and Grounding

This work formulates active learning for VG under the realistic setting where only raw images are available without accompanying text, and introduces Referred Region Ambiguity, a new acquisition function that measures whether the model's confidence collapses onto a single region or disperses across multiple candidates.

Junbeom Hong, Seonghoon Yu, Hyungsik Jung et al. · 0 citations
Aug 2026

Text-guided zero-shot localization of unseen object categories

A text-guided Zero-Shot Localization framework for unseen object categories (ZSOL) for addressing the aforementioned challenges, which can be guided by prompt words to identify and localize unseen object categories in images by transferring localization knowledge learned from supervised base categories.

Jingjing Wang, Xing-Lin Piao, Zongzhi Gao et al. · 0 citations
Conference Open access Sep 2026

F³S: feature fused few-shot segmentation with CLIP guided semantic priors and Sinkhorn attention refinement

Few-shot segmentation (FSS) remains a significant challenge due to the scarcity of annotated data and the need for precise object localization across novel classes. Existing approaches often rely on single-backbone architectures and coarse priors, which struggle to capture detailed semantics and precise spatial alignme...

Guo-Hua Geng, Xiao-Feng Wang · 0 citations
Sep 2026

Making Large Vision–Language Models Better Few-Shot Learners

Few-shot classification (FSC) aims to emulate the human ability to rapidly learn new concepts from a handful of examples. Large Vision-Language Models (LVLMs), with their rich prior knowledge and powerful visio-linguistic understanding capabilities, are emerging as a highly promising paradigm for FSC. This paper invest...

Chuan-Yi Zhang, Fan Liu, Yi Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.