Skip to content
Open access

A Unified Multimodal Search Framework Using Generative AI and Image Understanding for Enhanced Information Retrieval

Jul 2026 · Journal of Organizational and End User Computing · 0 citations

Abstract

Visual search systems often rely on image-only embeddings, limiting semantic understanding—especially with visually similar but semantically distinct items. To overcome this, the authors propose a unified multimodal search framework that integrates generative AI-based captioning and image understanding for improved retrieval. Using Qwen2-VL-7B, the system generates rich captions from images, then fuses visual and textual embeddings via cross-attention. A re-ranking module refines the results. Evaluated on COCO Captions and Fashion-Gen, the model achieves BLEU scores of 0.78 and 0.75, CIDEr scores of 1.12 and 1.08, and retrieval mAP of 0.75 and 0.73. It also records NDCG@10 of 0.91 and 0.89, outperforming baselines like FineCaption, SuperCap, and MM-Transformer by up to 9%. These results validate the approach in bridging the semantic gap between images and text, enabling more accurate, context-aware search in applications such as e-commerce, multimedia, and large-scale retrieval.

Read PDF