Multimodal large language models can expose task-relevant evidence through generation, but producing useful evidence does not by itself determine how it enters a retrieval embedding. We study whether the semantic organization of that evidence can also specify where representations are read. To address this question, we...
Zi-Rong Chen, Fu-Da Ye, En-Jun Du et al.· 0 citations
RePair is introduced, guided by three principles---Validity, Minimality, and Locality---which mines false positives bidirectionally, applies LLM-guided counterfactual editing, and trains with a local hard-pair contrastive objective.
Siyi Liu, Xiao-Rong Zhu, En-Jun Du et al.· 0 citations
Large vision-language models used as listwise rerankers must jointly process visual tokens from tens of candidates per query, making token pruning essential for practical deployment. Existing pruning methods retain tokens by attention saliency, yet we show that saliency is systematically misaligned with ranking contrib...
Siyi Liu, Han-Jun Yang, Chen-Chen Zhang et al.· 1 citation
EviRank is recast multimodal image re-ranking as a semantic constraint satisfaction problem and proposed, which parses any query - text-only, image-only, or composed - into a unified evidence package: typed criteria across six semantic slots, each labelled required, forbidden, or ignorable.
Enjun Du, Siyi Liu, Zi-Rong Chen et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.