Vision-language models (VLMs) achieve strong zero-shot transferability but remain vulnerable to target-domain shifts at inference time. Test-time adaptation (TTA) offers a practical remedy, yet most existing VLM-TTA methods follow a prediction-side adaptation paradigm. They use test samples to adjust logits, prototypes...
SVD-based pruning and quantization have recently emerged as a promising strategy for the ultra-efficient compression of large language models. In these methods, compression is performed in two stages: components are first truncated, and the remaining ones are subsequently quantized. Although this decoupled pipeline ben...
In this paper, we introduce layer-wise curriculum learning for efficient LLM compression. The proposed method facilitates the knowledge transfer from the teacher model to the student model, utilizing a curriculum learning approach that begins with easier optimization tasks and progressively tackles harder ones. In orde...
Donggeon Lee, Dooyeon Na, Seungmin Oh et al.· 0 citations
This work proposes Colla-Q, a bit-allocation framework to maintain balanced performance across experts through an activation-entropy-based bit-width allocation algorithm that encourages each expert to operate collaboratively in the quantized model, thereby improving the overall MoE performance and reducing the dependen...
OverRep is proposed, an Overcomplete Reparameterization framework for structured LLM pruning that temporarily overparameterizes the recovery module during training to absorb complex knowledge distilled from the original model.
This work addresses limitations in transfer learning for vision-language models through transformation-aware prompt conditioning and a re-calibrated contrastive loss, and treats same-class samples as positives rather than distinct instances, enabling the model to learn domain-specific features more effectively.