When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization
The stronger method is evaluated across 15 typologically diverse XL-Sum languages organised into three sets, beating the CE-only baseline on 10/15 languages; gains are most reliable where the two teachers contribute complementary signal, and weakest where they have saturated or jointly weak target-language coverage.