Multimodal large language models can expose task-relevant evidence through generation, but producing useful evidence does not by itself determine how it enters a retrieval embedding. We study whether the semantic organization of that evidence can also specify where representations are read. To address this question, we...
Zi-Rong Chen, Fu-Da Ye, En-Jun Du et al.· 0 citations
RePair is introduced, guided by three principles---Validity, Minimality, and Locality---which mines false positives bidirectionally, applies LLM-guided counterfactual editing, and trains with a local hard-pair contrastive objective.
Siyi Liu, Xiao-Rong Zhu, En-Jun Du et al.· 0 citations
Large vision-language models used as listwise rerankers must jointly process visual tokens from tens of candidates per query, making token pruning essential for practical deployment. Existing pruning methods retain tokens by attention saliency, yet we show that saliency is systematically misaligned with ranking contrib...
Siyi Liu, Han-Jun Yang, Chen-Chen Zhang et al.· 1 citation
Multimodal large language models (MLLMs) process hundreds or thousands of visual tokens per image, incurring prohibitive inference costs. While existing vision token pruning methods mitigate this overhead, they implicitly assume that a single fixed pruning strategy can be applied uniformly across all inputs. Our analys...
Hai-Jin Liang, P. Zhou, Zheng-Lin Wan et al.· 0 citations
Multi-objective ranking serves as the backbone of industrial information retrieval, requiring a holistic assessment of documents across dimensions such as Relevance, Authority, and Recency. The prevailing industry paradigm relies on ensembles of specialized BERT-based models, which are costly to maintain and fundamenta...
De-Zhi Ye, Junwei Hu, Xiaoyang Chen et al.· Proceedings of the 32nd ACM...· 0 citations
EviRank is recast multimodal image re-ranking as a semantic constraint satisfaction problem and proposed, which parses any query - text-only, image-only, or composed - into a unified evidence package: typed criteria across six semantic slots, each labelled required, forbidden, or ignorable.
Enjun Du, Siyi Liu, Zi-Rong Chen et al.· 2 citations
These findings suggest that, within the pointwise scoring paradigm, routing continuous relevance semantics through discrete text constrains ranking signal resolution reveals a bottleneck that is stable and difficult to overcome under current standard methods, rather than an easily resolvable training bias.
Xiaoyang Chen, Jie Liu, Haijin Liang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.