Skip to content
Preprint

Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval

Aug 2026 · 0 citations · 30 references
Computer Science

TL;DR

This study provides the first direct comparison of native multimodal embeddings against LLM-based visual ranking on Flickr30k, and observes that GPT-4.1 and Claude Sonnet 4.6 perform on par with Gemini Embedding 2.

Abstract

Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning. The March 2026 release of Gemini Embedding 2, Google's first natively multimodal embedding model to map text, images, video, audio, and documents into a single shared space, raises competition among multimodal retrieval systems. Simultaneously, frontier Large language models (LLMs) have also demonstrated strong visual understanding, raising the question of whether they can serve as effective zero-shot rankers. Our study provides the first direct comparison of native multimodal embeddings against LLM-based visual ranking on Flickr30k. We observe that GPT-4.1 and Claude Sonnet 4.6 perform on par with Gemini Embedding 2. Additionally, once embeddings are precomputed, multimodal embeddings are better suited for low-latency applications.

View source

Similar papers

Open access Jul 2026

A Unified Multimodal Search Framework Using Generative AI and Image Understanding for Enhanced Information Retrieval

A unified multimodal search framework that integrates generative AI-based captioning and image understanding for improved retrieval, enabling more accurate, context-aware search in applications such as e-commerce, multimedia, and large-scale retrieval.

Saeed Alzahrani, Farah Mohammad, Nazar Hussain · 0 citations
Preprint Aug 2026

Giga-Embeddings: Mixture-of-Experts Encoders for High-Throughput Text Embeddings

This work introduces Giga-Embeddings, a family of text embedding models designed to combine strong retrieval quality with efficient serving, and trains the compact model using a dimension-agnostic objective that aligns teacher and student similarity distributions.

Egor Kolodin, Egor Krasnoperov, Evgeniy Kosarev et al. · 0 citations
Preprint Sep 2026

UOT-Gap: A Variational Principle for the Modality Gap in Vision-Language Models via Unbalanced Optimal Transport

UOT-Gap is introduced, a training-free variational diagnostic that models frozen image and text embeddings with unbalanced entropic optimal transport (UOT) and establishes UOT-Gap as a diagnostic for caption quality, modality alignment, and retrieval robustness.

Zong-Lin Yang, Hui-Lan Ma, Xu-Dan Zheng et al. · 0 citations
Preprint Sep 2026

MARS: What Retrieval Signals Are Hidden in Multimodal Large Language Models for Text-Video Retrieval?

Text-video retrieval requires representations that can distinguish videos with similar scenes, actions, and temporal patterns. Recent multimodal large language models have been adapted as embedding models, but they often represent each input using a single token from the final layer. This can compress diverse video-tex...

Uicheol Jung, Juyoung Hong, Geuntaek Lim et al. · 0 citations
Preprint Aug 2026

WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report

WeMM-Embedding is presented, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs with flexible output dimensions and achieves leading performance on multiple public benchmarks.

Jun-Jie Zhou, Ke Mei, Lei Li et al. · 5 citations
Preprint Aug 2026

PUMA: Post-Hoc Sparsification of Universal Multimodal Embeddings for Efficient Retrieval

Universal multimodal embedders enable retrieval across text, image, and combined queries, but their dense representations incur high memory and inference costs. Post-hoc sparsification could reduce these costs but remains underexplored for multimodal retrieval. We introduce PUMA, a sparse autoencoder recipe that maps u...

Matteo Attimonelli, Alessandro De Bellis, F. M. Nardini et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.