AdaVSkip is proposed, which equips each layer with two lightweight routers that independently determine whether visual tokens pass through by or skip the self-attention and MLP modules, and maintains strong task performance with substantially less computation.
Yu-Yao Sun, Tao Deng, Shuang-Hua Li et al.· 0 citations
Middle-layer Attention Prediction (MAP), which uses Question Contrastive Teacher Selection to identify a sample-specific teacher layer by contrasting attention under the original and reference questions, and distills attention from the selected layer into a lightweight predictor that estimates visual token importance f...
Yu-Yao Sun, Tao Deng, Shuang-Hua Li et al.· 0 citations
PreGress is proposed, the first ranking-native pre-training and prompting framework for supporting a wide range of node ranking tasks, and designs lightweight, task-specific prompt modules that adapt a frozen ranking backbone to downstream tasks without full retraining.
Lujie Ban, Jiasheng Shi, Ying-Li Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.