This study provides the first direct comparison of native multimodal embeddings against LLM-based visual ranking on Flickr30k, and observes that GPT-4.1 and Claude Sonnet 4.6 perform on par with Gemini Embedding 2.
Abstract
Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning. The March 2026 release of Gemini Embedding 2, Google's first natively multimodal embedding model to map text, images, video, audio, and documents into a single shared space, raises competition among multimodal retrieval systems. Simultaneously, frontier Large language models (LLMs) have also demonstrated strong visual understanding, raising the question of whether they can serve as effective zero-shot rankers. Our study provides the first direct comparison of native multimodal embeddings against LLM-based visual ranking on Flickr30k. We observe that GPT-4.1 and Claude Sonnet 4.6 perform on par with Gemini Embedding 2. Additionally, once embeddings are precomputed, multimodal embeddings are better suited for low-latency applications.
A unified multimodal search framework that integrates generative AI-based captioning and image understanding for improved retrieval, enabling more accurate, context-aware search in applications such as e-commerce, multimedia, and large-scale retrieval.
Saeed Alzahrani, Farah Mohammad, Nazar Hussain· Journal of Organizational an...· 0 citations
This work introduces Giga-Embeddings, a family of text embedding models designed to combine strong retrieval quality with efficient serving, and trains the compact model using a dimension-agnostic objective that aligns teacher and student similarity distributions.
Egor Kolodin, Egor Krasnoperov, Evgeniy Kosarev et al.· 0 citations
UOT-Gap is introduced, a training-free variational diagnostic that models frozen image and text embeddings with unbalanced entropic optimal transport (UOT) and establishes UOT-Gap as a diagnostic for caption quality, modality alignment, and retrieval robustness.
Zong-Lin Yang, Hui-Lan Ma, Xu-Dan Zheng et al.· 0 citations
Text-video retrieval requires representations that can distinguish videos with similar scenes, actions, and temporal patterns. Recent multimodal large language models have been adapted as embedding models, but they often represent each input using a single token from the final layer. This can compress diverse video-tex...
Uicheol Jung, Juyoung Hong, Geuntaek Lim et al.· 0 citations
WeMM-Embedding is presented, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs with flexible output dimensions and achieves leading performance on multiple public benchmarks.
Universal multimodal embedders enable retrieval across text, image, and combined queries, but their dense representations incur high memory and inference costs. Post-hoc sparsification could reduce these costs but remains underexplored for multimodal retrieval. We introduce PUMA, a sparse autoencoder recipe that maps u...
Matteo Attimonelli, Alessandro De Bellis, F. M. Nardini et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.