This work introduces an optimal transport-based online clustering mechanism to automatically discover latent coarse-grained and fine-grained semantic priors without external supervision and introduces a cross-granularity alignment module that enforces consistency between these discovered latent structures and the target task, thereby regularizing the feature space against noise.
Abstract
While Semi-Supervised Learning (SSL) has witnessed substantial advancements, prevailing methods primarily operate at a single semantic granularity, overlooking the multi-scale semantic structures intrinsic to visual data. This limitation restricts the exploitation of rich feature constraints, often yielding representations that lack hierarchical coherence. To bridge this gap, we propose HCCL, a Hierarchical Consistency-driven Contrastive Learning framework. Unlike previous approaches that rely on fixed taxonomies, HCCL introduces an optimal transport-based online clustering mechanism to automatically discover latent coarse-grained and fine-grained semantic priors without external supervision. Specifically, we implement a resource-efficient prototype allocation strategy to capture salient intra-class variations and global semantic groupings. We then introduce a cross-granularity alignment module that enforces consistency between these discovered latent structures and the target task, thereby regularizing the feature space against noise. Furthermore, we construct adaptive affinity graphs based on these multi-granularity predictions to guide contrastive learning. Extensive experiments on CIFAR-10, CIFAR-100, STL-10, and Mini-ImageNet demonstrate that HCCL achieves highly competitive results against state-of-the-art competitors, particularly in challenging low-label regimes.
Evaluating semantic similarity between videos is a fundamental challenge in computer vision, essential for tasks ranging from out-of-distribution (OOD) detection to video retrieval. However, defining and labeling video similarity is notoriously difficult and expensive due to the complex spatio-temporal nature. In this...
Enrico Pallotta, Sina Raoufi, Lars Doorenbos et al.· 0 citations
This work presents a novel self-supervised architecture centered on graph prototype learning that sets a new state-of-the-art on the ARMM dataset with an accuracy of 95.70%, substantiating the efficacy and transferability of prototype-guided self-supervised learning for skeleton-based action representation.
Zhijie Xu, Hong-Wei Chen, Xia Li· International Journal of Mac...· 0 citations
Efficient Unsupervised Domain Adaptation (EUDA) is proposed, a parameter-efficient framework that leverages a frozen DINOv2 backbone as a feature extractor and updates only a lightweight bottleneck and classification head to promote both discriminative learning and cross-domain alignment.
Ali Abedi, Q. M. Jonathan Wu, Ning Zhang et al.· International Journal of Mac...· 9 citations
Unsupervised Domain Adaptation for Semantic Segmentation (UDA-SS) has seen significant progress in recent years. Existing UDA-SS approaches mostly adopt a pseudo-labeling schema to adapt model in the target domain, but they often overlook the inherent long-tailed data distribution in segmentation. We find that such sca...
Yi-Bo Wang, Rui-Kang Xu, Guang-Cheng Zhu et al.· Proceedings of the Thirty-Fi...· 0 citations
A Semantic Manifold-Aware Similarity Learning (SMSL) framework, where the term "semantic manifold" is used in an operational sense to denote a topology-aware organization of textual semantic units induced from token/entity interactions, rather than a strict low-dimensional differentiable manifold.
Li-Qi Zhu, Dezhi Han, Chong-Qing Chen· Neural Networks· 0 citations
The Dual-tOpology learning with adapTive Anchors (DOTA) is proposed, which not only learns the sample-anchor relationship but also preserves the topology structure among anchors, significantly enhancing the discriminability of learned representation while preserving the underlying data manifold.
Cheng-Long Zhang, Chao Zhang, Jun-Hao Zhang et al.· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.