SparseVLM+: Visual Token Sparsification With Improved Text-Visual Attention Pattern.
Improved text-visual attention patterns are introduced to enhance the fidelity of query-aware vision token selection and the Attention Gravity effect is correct, and a rank-based strategy to adaptively determine the sparsification ratio for each layer is introduced.