Frequency-Decoupled Cross-Attention Knowledge Distillation (FD-CanKD) is presented as a detector-oriented framework that transfers teacher knowledge at three complementary levels: head-level prediction supervision, relation-level non-local context transfer, and frequency-level component-selective alignment.
Abstract
Compact object detectors are suitable for resource-constrained visual perception, but their limited representation capacity creates an accuracy gap relative to large models. Conventional detector distillation often relies on prediction-level supervision or a single feature-alignment target, such as response, distribution, correlation, or frequency-domain matching. Frequency-Decoupled Cross-Attention Knowledge Distillation (FD-CanKD) is presented as a detector-oriented framework that transfers teacher knowledge at three complementary levels: head-level prediction supervision, relation-level non-local context transfer, and frequency-level component-selective alignment. Student features first aggregate teacher-side spatial context through cross-attention-based relation transfer, after which frequency-aware alignment preserves complementary structural and detail-sensitive cues. Under controlled Microsoft Common Objects in Context (COCO) experiments, fixed 50-epoch from-scratch comparisons show that FD-CanKD remains competitive with representative detector knowledge distillation baselines. Post-distillation continued fine-tuning further produces a stronger refinement-ready student than detector-only fine-tuning, reaching 48.87 mean average precision (mAP) at intersection-over-union thresholds from 0.50 to 0.95 (mAP50:95), 65.84 mAP50, and 53.40 mAP75 after 20 additional epochs. All distillation modules are removed after training, leaving the deployed student unchanged at 19.7M parameters. The framework is instantiated and evaluated in a controlled YOLOv12 teacher-student setting as a representative compact-detector case study.
This article introduces a novel object detection framework that integrates localization and adaptive spatial attention distillation techniques. While effective, prior knowledge distillation (KD) methods face a critical challenge in harmonizing feature-based and logit-based philosophies, often providing either entangled...
Zheng-Liang Lai, Jing-Ming Guo, Cheng-Ying Yang et al.· IEEE Canadian Journal of Ele...· 0 citations
Self-attention-based vision backbones perform well on dense prediction, but the quadratic computational cost of global softmax attention limits their efficiency as image resolution increases. We introduce Gated Token Recurrence (GTR), a softmax-free recurrent vision backbone that combines gated linear attention, altern...
Zhe Feng, Long-Fei Liu, Wei Liu et al.· 0 citations
Open-vocabulary aerial object detection aims to detect objects outside the training sets, which typically involves distilling knowledge from pretrained vision-language models, e.g., RemoteCLIP, to inherit its generalizable recognition ability and thereby enabling the models to detect novel categories. However, existing...
Shu Tian, Zhen-Dong Huang, Lin Cao et al.· IEEE Journal of Selected Top...· 0 citations
The results support cross-resolution KD as a practical way to recover fine-grained information during training while preserving low-resolution inference cost and show that a stronger teacher does not automatically produce a stronger student: transfer quality depends on the balance between teacher supervision strength a...
Chen-Qiang Li· Applied and Computational En...· 0 citations
Referring Expression Comprehension (REC) is commonly studied under dataset-specific fine-tuning, resulting in specialist models with limited cross-dataset generalization. In this work, we revisit REC from the perspective of unified open-vocabulary grounding and identify representation degeneration as a key obstacle to...
Jun-Yi Hu, Tian Bai, Feng-Yi Wu et al.· 0 citations
Experimental results show that the proposed SFDA object detection framework for the YOLO family of single-stage detectors consistently outperforms the baseline on multiple cross-domain detection tasks, demonstrating its effectiveness and good generalization ability under the source-free setting.