Progressive autoregressive image codecs provide an appealing paradigm for generative compression by quantizing continuous latents into discrete tokens, transmitting coarse-to-fine prefix tokens and generating the remaining suffix tokens at the decoder. However, their reconstruction quality is fundamentally limited by t...
Qing-Lei Yan, Rui-Xiao Dong, Yu-Tao Xie et al.· 0 citations
Black-box distillation is a practical route for transferring capabilities from API-accessible large language models that expose only text outputs into smaller student models. Recent on-policy adversarial methods such as GAD improve over SeqKD by forming an adversarial loop between a critic and a student, where the crit...
Knowledge distillation (KD) has become a prevalent technique for compressing large language models (LLMs). Existing KD methods are constrained by the need for identical tokenizers (i.e., vocabularies) between teacher and student models, as they assume a consistent semantic correspondence across logit dimensions, limiti...
Xiao Cui, Mo Zhu, Yu-Lei Qin et al.· IEEE Transactions on Pattern...· 0 citations