Multimodal graph-based fusion via image descriptions for few-shot open-set recognition.
Few-shot open-set recognition (FSOR) presents unique challenges due to the limited labeled samples and the presence of unknown classes during inference. While recent few-shot learning methods leverage class-level textual information to enhance feature discrimination, they often overlook image-grounded descriptions that...