Skip to content
Book Open access

Search Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses

Jul 2026 · International Conference on the Theory of Information Retrieval · pp. 403-413 · 0 citations · 35 references
Computer Science

TL;DR

AutoNuggetizer, a nugget-based framework, is employed to analyze ?

Abstract

Side-by-side comparisons that elicit human preferences are commonly used to assess the quality of large language model (LLM) output and have been applied to retrieval-augmented generation (RAG) systems. However, when applied to complex, information-seeking queries, these methods are limited by their inability to provide explanatory or diagnostic insights. As an alternative, nugget-based evaluations that decompose long-form answers into atomic facts have emerged as a promising strategy for RAG evaluation. In this work, we employ AutoNuggetizer, a nugget-based framework, to analyze ?5K Search Arena battles from LMArena by automatically generating and assigning nuggets, converting each model response into a quantitative score. We observe strong alignment between nugget-based Elo rankings and human preferences, exceeding the corresponding alignment achieved by LLM-as-a-judge with chain-of- thought (CoT) evaluation, while substantially reducing the number of preference inversions. Furthermore, we provide in-depth analyses including inversions, nugget quality, and shared-blindness effects. All our code is available at https://github.com/castorini/SxSNuggets.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

This work systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026, using staged screening and automated full-text coding to examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms.

Chao Wang · 0 citations
Preprint Aug 2026

Who's Keeping Score? Interactive Steering of LLM-Powered Scoring with Attune

Attune is presented, a mixed-initiative system for steerable LLM-powered scoring that performs pairwise comparisons across records to develop a global understanding first, and then resolves these comparisons into consistent score assignments-deriving scoring criteria and rules bottom-up in the process.

Bhavya Chopra, Meng Chen, Rebecca Dang et al. · 0 citations
#natural language process... Preprint Sep 2026

LLJ Cards: Best practices for the Use of LLMs as Judges

In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these systems have been widely adopted by researchers and practitioners across a broad range of measurement tasks, driven by their strong performance, scalability, and cost-eff...

Khaoula Chehbouni, Melina Medjdoub, Florian Carichon et al. · 0 citations
Review Open access Aug 2026

Large language models as judges: recent advances in LLM-based evaluation, critique, preference modeling, and feedback for text and code

This survey provides a comprehensive overview of recent advances in LLM-based evaluation, covering techniques, applications, and challenges across domains, with future directions emphasizing standardized protocols, uncertainty estimation, and human–AI collaboration.

M. Nadăş · 0 citations

Understanding the gap in evaluation frameworks and metrics for RAG-based question-answering system

A systematic analysis of existing evaluation frameworks and metrics for RAG-based Question Answering (QA) systems reveals urgent needs for hybrid evaluation frameworks, reasoning path traceability, combined pipeline assessment, and harmlessness evaluation, laying groundwork for the next generation of evaluation methodo...

Nakul Mehta, A. Ojo, Edward Curry · 0 citations
#natural language process... Preprint Sep 2026

On Epistemic Diversity in Large Language Models

It is argued that epistemic diversity is a useful evaluation dimension for settings where LLMs are used to support knowledge-intensive tasks, and a preliminary framework for conceptualizing and measuring epistemic diversity in LLMs is proposed, and operationalized in two domains.

Elisabeth Kirsten, N. Krämer, Muhammad Bilal Zafar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.