The proposed D3O, a dynamic distribution distillation framework that replaces static supervision with training-driven evolution of ordinal label distributions via self-distillation, introduces a contrastive ordinal-aware label enhancement module that leverages vision-language alignment to recover refined label distributions capturing both inter-class ambiguity and instance-level uncertainty.
Abstract
Ordinal regression is widely used in scenarios where labels are discrete yet inherently ordered. In practice, however, ordinal labels are often obtained by discretizing underlying continuous semantics through subjective human judgment, resulting in ambiguous boundaries and annotation noise. Such uncertainty challenges existing methods that rely on fixed supervision targets, which may reinforce biased ordering under subjective annotations. To address this limitation, we propose D3O, a dynamic distribution distillation framework that replaces static supervision with training-driven evolution of ordinal label distributions via self-distillation. Specifically, we introduce a contrastive ordinal-aware label enhancement module that leverages vision-language alignment to recover refined label distributions capturing both inter-class ambiguity and instance-level uncertainty. Furthermore, we design a CDF-based cross-layer interaction distillation mechanism to propagate cumulative ordinal structure across network hierarchy, ensuring consistent ordinal geometry in intermediate representations. Extensive experiments on four general ordinal regression tasks demonstrate that our proposed D3O consistently outperforms existing approaches, particularly under severe class imbalance and noisy supervision. These results highlight the effectiveness of dynamic supervision in learning robust ordinal representations beyond fixed targets. The code will be publicly available.
Experiments on two multi-expert ulcerative colitis endoscopic-image datasets under two ordinal-noise models show that Ord-NLL is competitive with or superior to strong baselines while reducing mean absolute error, and that Ord-NLL+ often yields further gains.
Shumpei Takezaki, K. Shiku, S. Harada et al.· IEEE Access· 0 citations
This work proposes an integrated learning paradigm that simultaneously enhances feature compactness and improves robustness against label noise and introduces a feature disentanglement mechanism that isolates reliable label-related feature representations from spurious ones introduced by noisy supervision.
Yuzhi Tao, Anhui Tan· Computers, Materials & C...· 0 citations
The efficacy of logit transfer in knowledge distillation (KD) is often hampered by disparities in model architectures and imbalances in data distributions. Existing decoupled distillation methods predominantly rely on rigid category partitioning (e.g., target vs. non-target), failing to adapt to varying logit scales across different teacher models and showing instability in long-tailed learning scenarios. In this paper, we propose Head-Decoupled Logit Distillation (HDLD), a novel framework that redefines the distillation decoupling process from the perspective of instance-specific dynamic semantic structures. By introducing a dynamic gating mechanism, HDLD decouples logits into Head Category Knowledge Distillation (HCKD) and Non-Head Knowledge Distillation (NHKD). This approach adaptively extracts core semantic correlations based on sample characteristics, effectively filtering background noise in long-tailed settings and rectifying decision boundary collapse. Extensive experiments on image classification and object detection benchmarks demonstrate that HDLD consistently outperforms state-of-the-art methods, validating its superior robustness and transferability across diverse architectures and tasks. Our code can be found at https://github.com/Hans-KnowledgeDistillation/HDLD
Zheng-Han Ye, W. Lyu, Qing Guo et al.· IEEE Transactions on Image P...· 0 citations
This work introduces a novel framework, Gaussian Bridge Consistency (GBC), to address challenges of semi-supervised learning by constructing semantic interpolation paths between unlabeled samples and high-quality class anchors, and proposes BridgeMix, a confidence-aware feature mixing strategy that interpolates both sample and anchor pairs to amplify cross-sample generalization.
Hong-Yang He, Xin-Yuan Song, Yan Zhong et al.· 0 citations
This work proposes a teacher-student semi-supervised learning framework that generates high-quality pseudo-labels from unlabeled data through confidence-aware map refinement, and introduces a spatial clipping technique that selectively preserves high-confidence regions while removing unreliable segments.
Chikao Tsuchiya, Dhaval Bhanderi, David Ilstrup et al.· 0 citations
This work proposes Expert-Guided Mutual Distillation (EGMD), which learns what evidence to trust across the prediction pipeline, and constructs Weibo_Balanced, a domain-balanced benchmark that isolates the effect of imbalance on generalization.
Xuan Feng, Guihong Liu, Tianlong Gu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.