This work proposes Visual Contrastive Self-Distillation, namely VCSD, which converts image-content removal into an on-policy self-distillation signal, and consistently outperforms matched OPSD across Qwen3-VL and Qwen3.5 models.
Abstract
On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods create this asymmetry either through privileged answers or visual evidence. We ask whether both can be removed, yielding a simpler form of OPSD driven purely by input conditioning. For this purpose, we propose Visual Contrastive Self-Distillation, namely VCSD, which converts image-content removal into an on-policy self-distillation signal. At each student-generated response prefix, the EMA teacher produces two next-token distributions under the same prompt and prefix -- one conditioned on the original image and the other on a content-erased control. Their token-wise log-probability difference highlights candidates whose likelihood is specifically increased by the instance-level visual content. We use this contrast to sharpen the teacher's original-image distribution within its plausible support, and distill the resulting full-distribution target into the student. Using ViRL39K dataset, VCSD consistently outperforms matched OPSD across Qwen3-VL and Qwen3.5 models. For example, on Qwen3-VL, it improves the seven-benchmark aggregate from $62.27\% \rightarrow 67.04\%$ at 2B, $71.30\% \rightarrow 73.16\%$ at 4B, and $72.51\% \rightarrow 76.26\%$ at 8B. Furthermore, VCSD requires no external teacher, privileged answers, visual evidence signals, reasoning traces, or additional inference-time cost.
Self-Supervised Visual On-Policy Distillation (S$^2$VOPD), a simple yet effective method that constructs on-policy learning signals from asymmetric augmented views, systematically explores a broad design space of visual augmentations and uncover that asymmetry matters.
Yijiang Li, Yijun Liang, Yunjie Tian et al.· 0 citations
OPD-V is introduced, a visual OPSD paradigm that instantiates privileged information through the Positive Teacher and Negative Teacher that consistently improves reasoning performance while reducing training cost.
This work introduces CVPD (Contrastive Counterfactual Visual Process Distillation), which is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs, and proposes a three-gate Counterfactual Criterion that identifies visual blind spots where zooming into a region changes and sharpens the model's answer distribution.
Shravan Venkatraman, Omkar Thawakar, Ritesh Thawkar et al.· 1 citation
Fisher-Projected On-Policy Distillation (FP-OPD), which distills only locally realizable teacher corrections provide a more effective target for distilling compact vision--language models.
Leyan Xue, Feng Xiong, Ming-Jun Ma et al.· 3 citations
It is demonstrated that resolution differences can serve as a simple and scalable source of privileged information, providing an effective and efficient approach to on-policy self-distillation for multimodal large language models.
U-OPSD first samples multiple rollouts and constructs a pseudo solution by majority vote under a self-consistency threshold, and conditions the model's distribution on the pseudo-solution and distills itself on the disagreeing completions, allowing the model to correct itself precisely where it is confidently wrong.
Yijiang Li, Bingyang Wang, Yijun Liang et al.· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.