The efficacy of logit transfer in knowledge distillation (KD) is often hampered by disparities in model architectures and imbalances in data distributions. Existing decoupled distillation methods predominantly rely on rigid category partitioning (e.g., target vs. non-target), failing to adapt to varying logit scales across different teacher models and showing instability in long-tailed learning scenarios. In this paper, we propose Head-Decoupled Logit Distillation (HDLD), a novel framework that redefines the distillation decoupling process from the perspective of instance-specific dynamic semantic structures. By introducing a dynamic gating mechanism, HDLD decouples logits into Head Category Knowledge Distillation (HCKD) and Non-Head Knowledge Distillation (NHKD). This approach adaptively extracts core semantic correlations based on sample characteristics, effectively filtering background noise in long-tailed settings and rectifying decision boundary collapse. Extensive experiments on image classification and object detection benchmarks demonstrate that HDLD consistently outperforms state-of-the-art methods, validating its superior robustness and transferability across diverse architectures and tasks. Our code can be found at https://github.com/Hans-KnowledgeDistillation/HDLD
Zheng-Han Ye, W. Lyu, Qing Guo et al.· IEEE Transactions on Image P...· 0 citations
Remote sensing object detection (RSOD) faces significant challenges due to complex background clutter and the loss of fine-grained features in deep networks. To address these issues, we propose a novel dynamic feature-focused detection transformer, termed DynaFocus-DETR. First, we introduce a Dynamic Feature-Focused Transformer (DynaFocus-Transformer) module, which leverages the self-attention weights of high-level features to dynamically extract and fuse local details from lower-level features, thereby suppressing background interference and enhancing semantic alignment. Second, we design a dual-branch context-aware downsampling (DBCAD) module to reduce information loss during downsampling by fusing max-pooling features with contextually enriched features extracted via adaptive kernels. Finally, we design a Density Map-Guided Query Selection (DMGQS) method to provide high-quality queries for the decoder of the transformer. Extensive experiments on the DIOR and NWPU VHR-10 datasets demonstrate the superiority of our approach, achieving state-of-the-art mAPs of 77.9% and 94.0%, respectively. Code is available at https://github.com/Zhang-Haoyan/DynaFocus-DETR
Haoyan Zhang, W. Lyu, Qing Guo et al.· IEEE Geoscience and Remote S...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.