Skip to content
Preprint

Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search

Aug 2026 · 0 citations · 27 references
Computer Science

TL;DR

A VLM-based automated relevance evaluation pipeline deployed within Pinterest Search for online A/B experiments is presented, demonstrating that VLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency.

Abstract

Relevance evaluation plays a crucial role in personalized search systems, serving as a guardrail alongside user engagement metrics to ensure that search results align with user queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present a VLM-based automated relevance evaluation pipeline deployed within Pinterest Search for online A/B experiments. We rigorously validate the alignment between VLM-generated judgments and human annotations, demonstrating that VLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency. Leveraging VLM-based labeling further unlocks opportunities to expand the query set, optimize sampling design, and efficiently assess a wider range of search experiences at scale. This approach leads to higher-quality relevance metrics and significantly reduces the Minimum Detectable Effects (MDEs) in online experiment measurements.

View source

Similar papers

Book Open access Aug 2026

AutoDavis: Automatic and Dynamic Evaluation Protocol of Large Vision-Language Models on Visual Question-Answering

AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.

Han Bao, Yue Huang, Yan-Bo Wang et al. · 0 citations
Conference Open access Sep 2025

DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values

DiverValue-Bench is introduced, a population-aware benchmark for evaluating multi-dimensional value alignment across 74 countries/regions and it is shown that lightweight preference-based fine-tuning with Low-Rank Adaptation and Direct Preference Optimization substantially improves in-domain value alignment while yield...

Yao Liang, Dongcheng Zhao, Fei-Fei Zhao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities

This work introduces AgenticInterleave, a single-agent ReAct framework for retrieval-supported answer generation, together with IVR-12, a 12-dimensional rubric for assessing the content, presentation, and image quality of interleaved references and model outputs, and evaluates both short-answer correctness and interlea...

Hao-Nan Jiang, Guo-Jian Zhan, Jian-Cong Xie et al. · 0 citations
Book Open access Aug 2026

Automating End-to-End Hybrid Query Processing: Benchmark, Solution, and Insights

A large-scale benchmark with 60\sim 90× more queries than prior work, built on 3× more databases, an automated pipeline that can execute existing methods without manual intervention, and multi-dimensional, fine-grained evaluation metrics for comprehensive assessment.

Bo Li, Chenzhan Wang, Long-Kang Lin et al. · 0 citations
Preprint Aug 2026

From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation

A sequential behavioral alignment framework pairing fine-tuning with preference optimization over paired correct and counterfactual rationales is developed and applied, demonstrating that behavioral alignment mitigates bidirectional rationalization while delivering human-interpretable reasoning traces without manual pi...

Alireza S. Ziabari, Kat Ellis, Colleen E. Chan et al. · 0 citations
#natural language process... Preprint Aug 2026

Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

This work proposes prediction-powered evaluation, a framework that combines limited human judgments with large-scale automatic scores to obtain data-efficient system comparisons that are provably unbiased, and introduces the Prediction-Powered Saving Ratio (PPSR), a meta-metric that measures how much human annotation a...

Mingqi Gao, Anthony B. Sicilia, Weiye Shi · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.