Skip to content
Preprint

IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves

Aug 2026 · 0 citations · 77 references
Computer Science

TL;DR

It is shown that information contained in the teachers can be leveraged to tailor the noise for multi-teacher distillation, and a method is proposed that generates teacher-specific, improved samples optimized for data-free distillation.

Abstract

Multi-teacher distillation has emerged as a way to combine complementary teacher models into a single student model that exhibits the strengths of all its teachers. The student is trained to mimic the output of the teachers on a set of images, typically the union of the individual teacher's training sets, assuming this data is available. In this paper, we question that assumption and explore alternative options. We first study how far one can go when distilling from teachers fed with different types of noise. Then, we show that information contained in the teachers can be leveraged to tailor the noise for multi-teacher distillation: we propose a method that, thanks to decorrelation losses at both patch and image levels, generates teacher-specific, improved samples optimized for data-free distillation. Experiments show that our most effective samples, IDeaL, lead to strong students that successfully capture complementary information from the teachers, yielding surprisingly competitive results that substantially narrow the gap with students distilled from real images. Moreover, given a limited budget of 1K images for distillation, students distilled using our IDeaL samples match or surpass the performance of those distilled using a 1K-image subset of ImageNet.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Knowledge Distillation under Teacher Misspecification: An Order-Parameter Analysis of the Gap between Teacher Mimicry and Task Performance

It is proved that the learning dynamics and the distillation error $\Ets$ are exactly invariant to $\dmiss$, whereas the true error $\Etzs$ and the gap $\Delta=\Etzs-\Ets$ are strictly increasing in $\dmiss$, with a rate that is amplified linearly by the complexity $M_0$ of the true teacher.

K. Hara, H. Hino · 0 citations
Preprint Aug 2026

Too much of a good thing -- when knowledge distillation promotes overfitting, and how to avoid it

This work proposes a student design based on simple, homogeneous blocks mirroring those of the teacher, distilling knowledge between corresponding blocks, showing that intermediate block-wise distillation, guided appropriately, is key to building compact data-efficient models without sacrificing accuracy.

Irene Trigueros-Lorca, Leonardo Concepción, Christian Wagner et al. · 0 citations
Aug 2026

Making Knowledge Distillation Open Again

Knowledge distillation (KD) has become a pivotal technique for transferring knowledge from large-scale teacher models to lightweight student models. However, traditional feature-based distillation methods necessitate the direct exposure of the teacher’s intermediate representations, raising concerns regarding data priv...

Jun-Fei Yi, Si-Hao Lin, Hui Zhang et al. · 0 citations
Preprint Aug 2026

Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable

Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effect: larger datasets can make subtle teacher-specific signals easier to detect in the trained student, even when examples are off-task and ne...

Zhichen Dong, Zhi-Xuan Liu, Yuanjiu Fan, et al. · 1 citation
Open access Jul 2026

Knowledge Distillation Across Tasks and Model Families: A Comparative Simulation Study—Denoising, Dark Knowledge, and the Role of Model Capacity

Knowledge distillation transfers information from a high-capacity teacher to a smaller student, but its behavior across regression, classification, and heterogeneous tabular model families remains insufficiently understood. This paper presents a comparative simulation study of distillation in structured-data settings,...

Bogdan Oancea · 0 citations
Preprint Aug 2026

SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features

SQuaT (Student-Aware Quantized Teacher Features), a label-free QAT framework with KD that theoretically eliminates this lower bound on the distillation loss by applying the student's quantization parameters to quantize the teacher's features during distillation is proposed.

H. Lee, Hyeonsik Jo, Jinwook Chung et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.