This work proposes prediction-powered evaluation, a framework that combines limited human judgments with large-scale automatic scores to obtain data-efficient system comparisons that are provably unbiased, and introduces the Prediction-Powered Saving Ratio (PPSR), a meta-metric that measures how much human annotation an automatic metric can save when used within prediction-powered evaluation.
Abstract
Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inference (PPI), we propose prediction-powered evaluation, a framework that combines limited human judgments with large-scale automatic scores to obtain data-efficient system comparisons that are provably unbiased. We develop parametric and non-parametric procedures, analyze the efficiency trade-off between paired and unpaired designs, and validate the framework on six WMT datasets. We further introduce the Prediction-Powered Saving Ratio (PPSR), a meta-metric that measures how much human annotation an automatic metric can save when used within prediction-powered evaluation. PPSR directly targets metric utility for prediction-powered evaluation and yields more discriminative and stable metric rankings than existing system-level meta-metrics. Overall, our new paradigm reframes automatic metrics as tools for reducing human annotation cost rather than replacing human judgment, and applies broadly to non-verifiable tasks.
This work introduces QUORUM (QUality-Optimized Routing Using Multiple annotators), a budget-aware routing framework that dynamically assigns each instance to human or LLM annotators under a fixed annotation budget and supports multiple annotations per instance, combining them through agreement-based rewards to improve...
Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu et al.· 1 citation
Human-model alignment is critical for trustworthy AI-assisted decision-making systems. Yet, most work evaluates model predictions against single ground-truth labels, overlooking that humans themselves often disagree on labels, a signal of genuine ambiguity. We investigate whether models struggle on the same instances t...
Human pairwise comparisons provide a reference for evaluating large language models (LLMs), but collecting sufficient judgments for each new release is costly and time-consuming. LLM judges offer a scalable alternative, although their comparisons may differ systematically from human preferences and across judges. We st...
Xin Zhou, Si-Nian Zhang, Zhan-Yan Yang et al.· 0 citations
Prompt sensitivity is widely treated as a model robustness deficiency. Yet the extent to which measured sensitivity reflects genuine model instability, rather than artifacts of the evaluation method used to measure it, remains largely underexplored. We introduce Evaluation-Attributable Sensitivity (EAS), a per-instance...
Sayumi Muthukumarana, Buddhi Wijenayake, R. Godaliyadda et al.· Moratuwa Engineering Researc...· 0 citations
NAPHA (eNtropy-Aware Post-Hoc Alignment), a simple yet effective lightweight post-hoc alignment method that matches the LLM distribution to the HJD by first assigning an instance to a discrete entropy class and then routing it to specialized, trained alignment models is proposed.
Sebastian Steindl, Nikos Voskarides, Alberto Gasparin et al.· 0 citations
Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question: does WER actually track how humans judge ASR transcript quality? We introduce HATS-en, an...
Hritika Sharma, Thibault Bañeras-Roux, Alessandra Pinto et al.· 0 citations
The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.
Large language models are increasingly integrated into systems that can retrieve information, call APIs, search databases, send messages, summarize documents, and take actions on behalf of users. These capabilities make LLMs useful, but they also increase the potential impact of LLM-related attacks. The post Prompt Injection Attacks in LLM Systems appeared first on GPT-Lab.