Streamlined Knowledge Distillation is proposed, a simple yet effective logit-based method that transfers only two essential forms of knowledge without requiring additional alignment or relational modeling and introduces a Mahalanobis distance-based direction-wise loss stabilized through Tikhonov regularization and Cholesky decomposition.
Knowledge distillation (KD) has become a pivotal technique for transferring knowledge from large-scale teacher models to lightweight student models. However, traditional feature-based distillation methods necessitate the direct exposure of the teacher’s intermediate representations, raising concerns regarding data priv...
Jun-Fei Yi, Si-Hao Lin, Hui Zhang et al.· IEEE Transactions on Image P...· 0 citations
Logit-based knowledge distillation (KD) is pivotal for efficient model compression and cross-architecture learning. However, conventional methods typically rely on static, single-scale logit alignment, thereby overlooking the semantic evolution trajectory embedded in cross-scale prediction transitions. To bridge this g...
Hejie Lu· Proceedings of the 32nd ACM...· 0 citations
Token-level knowledge distillation (KD) matches two conditional distributions per position, yet the standard objectives compare them pointwise: a Kullback-Leibler gradient is blind to which wrong token receives probability mass. We develop a distributional view in which the teacher is represented not by a single soften...
This work revisits on-policy Reverse Kullback-Leibler distillation and decomposes its objective into a teacher-fitting term and a student-entropy term, without introducing an explicit FKL branch, and proposes Adaptive Entropy Distillation (AED), which uses the teacher's entropy to dynamically calibrate token-level imit...
Shizheng Li, Zhiyu Shen, Yu-Yin Lu et al.· 0 citations
A practitioner's study of how to make distillation training efficient is presented, organised around two systems contributions, and a fused, chunked KL loss is introduced, making peak memory linear in the sequence length.
Bakbergen Ryskulov, Iker García-Ferrero, David Montero et al.· 0 citations
The stronger method is evaluated across 15 typologically diverse XL-Sum languages organised into three sets, beating the CE-only baseline on 10/15 languages; gains are most reliable where the two teachers contribute complementary signal, and weakest where they have saturated or jointly weak target-language coverage.
Dipto Sumit, Ankan Kumar Roy Srizon, Sadia Khair Rodela et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.