Skip to content
Preprint

MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

Aug 2026 · 0 citations · 49 references
Computer Science

Abstract

Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations. While unlabeled multi-modal data is abundant, it remains elusive how to exploit them for ICL. We propose MAG (MAnifold-Guided semi-supervised in-context demonstra- tion selection), an efficient framework that leverages unlabeled data to improve multi-modal ICL. MAG formulates demonstration selection as a semi-supervised propagation problem on a multi-modal graph and adopts a two-stage strategy: (i) relevance score propagation identifies a compact set of high-impact unlabeled samples for pseudo-labeling, reducing MLLM inference cost; (ii) multi-modal relevance is used to select the final demonstrations. We show that textual represen- tations are more effective for relevance propagation, while both visual and textual modalities are crucial for high-quality demonstration selection. Experiments on eight multi-modal benchmarks demonstrate that MAG consistently outperforms strong baselines in label-scarce regimes, achieving significant gains with a limited pseudo-labeling budget.

View source

Similar papers

Sep 2026

Multimodal graph-based fusion via image descriptions for few-shot open-set recognition.

A multimodal Graph-based Fusion (MGF) framework that learns visually grounded semantic representations to enhance FSOR performance and achieves superior open-set recognition and competitive closed-set classification performance.

Xilang Huang, Seon-Han Choi · 0 citations
#artificial intelligence Preprint Sep 2026

Follow the Geometry, Not the Model: Cold Start Semi-Supervised Learning

Modern semi-supervised learning (SSL) couples pseudo-label generation and classifier training, using the classifier's own confidence to select the pseudo-labels that are then used to update the model. In the cold-start regime, where at most a few labels per class are available, this coupling is ill-posed, since the cla...

Itai David, D. Weinshall · 0 citations
Open access Aug 2026

Weakly-Supervised Learning with Partial Multi-Labels by Leveraging Dual Label Correlation Perspectives

A novel PML method, namely Wasserstein Partial Multi-Label Learning with dual Label Correlation Perspectives (Wpml3cp), solved by the gradient descent with an augmented Lagrange multiplier technique, and empirical results demonstrate that Wpml3cp and Wpml3cp-D can outperform the PML baselines in various noisy levels.

Xi-Ming Li, Yuanchao Dai, Bing Wang et al. · 0 citations
Preprint Aug 2026

D3ER: Supporting Multi-Modal Recommendation via Disentangle and Distillation-based Dynamic Ensemble

Incorporating items'information shared among multiple modalities into a fused representation, multi-modal recommendation (MR) has demonstrated documented success than canonical unimodal recommendation. Although several attempts have been made to extract the discriminative information unique in each modality, existing m...

Bing-Nan Wang, Yi Li, Xiong-Xin Tang et al. · 0 citations
Aug 2026

Deep Semi-Supervised Learning via Tensor Label Propagation for High-Dimension-Low-Sample-Size Data.

Semi-supervised learning (SSL) aims to effectively utilize a small amount of labeled data together with a large volume of unlabeled data to improve learning performance. Among various SSL strategies, label propagation has been widely adopted due to its ability to diffuse label information across data points via graph s...

Hongmin Cai, Jiali Sun, Fei Qi et al. · 0 citations
Conference Aug 2026

Adaptive Cross-Modal Fusion With Instance-Level Gating for Vision-Language Understanding

Multimodal deep learning integrates heterogeneous data sources such as images and text to enable machines to understand complex real-world contexts. Although recent vision-language models have achieved significant progress, most existing approaches rely on rigid fusion strategies that combine modalities either at early...

Unnati A. Patel, Sanskruti Patel, J. Nanavati et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.