Generalized zero-shot learning (GZSL) aims to recognize novel categories by leveraging semantic knowledge transferred from previously observed categories, where learning transferable features for effective visual-semantic alignment plays a critical role. Conventional methods typically utilize all shared attributes to learn semantically related features. However, not all attributes are related to a specific instance, and the inconsistency may result in less effective image features. Additionally, the visual features extracted by the pre-trained backbone may be unsuitable for a particular sample, reducing the discrimination. To address these problems, we propose a novel instance-level visual and semantic adaptation framework that effectively adapts pre-trained image features for the GZSL tasks. From the semantic adaptation aspect, we propose to choose instancespecific attributes for dynamic semantic prompt tuning. From the visual adaptation aspect, we construct instance visual proto-types to produce channel attention, which adaptively strengthens crucial visual features for each sample. Extensive experiments on three GZSL benchmark datasets demonstrate that our approach achieves the new state-of-the-art performance.
Huajie Jiang, Zheng-Xian Li, Xi-Chao Yu et al.· IEEE Transactions on Image P...· 0 citations
Generalized zero-shot learning (GZSL) addresses the challenging task of recognizing both seen and unseen classes by leveraging shared semantic knowledge. A core challenge in this domain is achieving robust visual-semantic alignment to transfer knowledge from seen classes to novel classes. Current state-of-the-art methods typically fine-tune large-scale visual backbones on scarce training data. However, this approach frequently leads to severe overfitting to seen classes, which significantly degrades performance on novel categories. To mitigate this issue, we propose the Dual Adaptive Visual-Semantic Prompt Collaboration Network (VSPCN+), a novel framework that utilizes prompt-tuning for effective feature adaptation. Our method introduces a dual-prompt mechanism comprising both visual and semantic prompts. The semantic prompts guide the visual encoder to learn visual features that are more semantically consistent with class attributes, while the visual prompts steer the semantic encoder to generate semantic representations that are more visually grounded. This collaborative process enhances the overall visual-semantic consistency. A key innovation of our work is the dynamic generation of instance-adaptive prompts, which contrasts with existing prompt-learning methods that rely on static, global prompts. By tailoring prompts to individual instances, our approach enhances the model’s robustness and generalization capabilities across diverse visual inputs. This collaborative adaptation, guided by our dual-prompt mechanism, allows the visual and semantic encoders to produce consistent representations for effective visual-semantic alignment. Extensive experiments on standard GZSL benchmarks demonstrate that our proposed VSPCN+ performs favorably against several state-of-the-art methods.
Huajie Jiang, Zheng-Xian Li, Yuankai Qi et al.· International Journal of Com...· 0 citations
This benchmark provides practical guidance for model selection and argues that future models should be judged by biological generalisation, interpretability and perturbation-grounded validity, not by scale or leaderboard performance alone.
S. Chen, Roxana Zahedi, Lucy Chhuo et al.· 0 citations
A systematic review of the literature on methodologies, frameworks, and techniques for auditing AI systems, focusing on legal and ethical considerations and compliance with regulations, reveals gaps in current auditing practices and highlights the importance of incorporating AI value chain stages and AI maturity levels into auditing frameworks.
Usman Shahbaz, A. Beheshti, B. Abedin et al.· ACM Computing Surveys· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.