AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference
AdaVSkip is proposed, which equips each layer with two lightweight routers that independently determine whether visual tokens pass through by or skip the self-attention and MLP modules, and maintains strong task performance with substantially less computation.