MGTE: A Modular Multi-Granularity Text Ensemble for Interactive Image Retrieval
Interactive text-to-image retrieval seeks to identify a target image through multi-turn dialogue, where users progressively refine ambiguous search intent with additional visual constraints. Existing approaches commonly encode the full dialogue as a single concatenated sequence, but this strategy is vulnerable to token...