Skip to content
Conference

MGTE: A Modular Multi-Granularity Text Ensemble for Interactive Image Retrieval

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 658-663 · 0 citations · 15 references

Abstract

Interactive text-to-image retrieval seeks to identify a target image through multi-turn dialogue, where users progressively refine ambiguous search intent with additional visual constraints. Existing approaches commonly encode the full dialogue as a single concatenated sequence, but this strategy is vulnerable to token-budget truncation and may overlook late-turn refinements or explicitly negated attributes. We propose Multi-Granularity Text Ensemble (MGTE), a training-free scoring module that represents dialogue context from three complementary views: the full dialogue, the latest turn, and the initial caption. These views are combined with turn-adaptive weights to preserve accumulated context, emphasize recent refinements, and reduce semantic drift across turns. Since MGTE operates only at the text-scoring stage, it can be directly integrated into existing interactive retrieval pipelines without model retraining. For dialogues containing exclusion constraints, we further introduce Negative-Aware Scoring (NAS), which penalizes candidates that match negated concepts. Experiments on VisDial demonstrate that MGTE consistently improves three representative retrieval pipelines, achieving Hits@10 gains of up to +4.31. When combined with NAS, the DAR pipeline reaches 85.42 Hits@10. Further analysis shows that MGTE improves robustness against late-dialogue truncation and produces more stable final-turn rankings than latest-turn-only scoring.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.