Recent multimodal large language models (MLLMs) advance visual reasoning by strengthening both perception and reasoning, implicitly assuming a process that transitions seamlessly from perception to reasoning. However, we observe a counterintuitive phenomenon that challenges this assumption: holding the model, question,...
Jin-Tao Tong, Yujing Lou, Zhan-Ming Shen et al.· 0 citations
Large Vision-Language Models (LVLMs) incur high computational costs from redundant visual tokens. Although training-free attention-based multi-layer pruning in the vision encoder stage has been explored as an effective strategy, we find that pruning in shallow layers consistently degrades performance. In this paper, we...
Shuo Zhang, Jin-Tao Tong, Yixiong Zou et al.· 0 citations
Few-shot class-incremental learning (FSCIL) aims to incrementally learn novel classes with only a few samples while avoiding forgetting base classes. However, current methods show a tendency to misclassify novel-class samples into base classes, which we find to be caused by the excessive focus on base-class-discriminat...
Haichen Zhou, Y. Lyu, Yixiong Zou et al.· IEEE transactions on multime...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.