Multi-vector vision-language retrievers enable fine-grained Visual Document Retrieval (VDR) through late interaction, but storing and scoring hundreds of visual patch embeddings per page incurs substantial overhead. Existing training-free methods rely on pruning or merging: pruning degrades sharply under aggressive com...
A systematic analysis of expert routing patterns in MoE models reveals Language Routing Isolation, in which high- and low-resource languages tend to activate largely disjoint expert sets, and proposes RISE, a framework that exploits routing isolation to identify and adapt language-specific expert subnetworks.
Kening Zheng, Wei-Chieh Huang, Jiahao Huo et al.· arXiv.org· 4 citations· ⚡2
This work starts from an empirical observation: when query-relevant visual evidence is explicitly strengthened using the model's own attention, generation becomes more accurate, suggesting that many failures do not arise solely from missing perception, but from an insufficient tendency to trust the evidence the model h...
Xin Zou, Hao Deng, Yibo Yan et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.