Skip to content

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

Jul 2026 · arXiv.org · Vol abs/2607.28801 · 1 citation
Computer Science

TL;DR

A dataset-centric meta-evaluation framework that audits benchmark datasets at the sample level along five latent dimensions, providing a principled methodology for analyzing and re-composing existing benchmarks to better reflect diverse evaluation needs.

Abstract

Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of individual samples. We introduce a dataset-centric meta-evaluation framework that audits benchmark datasets at the sample level along five latent dimensions: 1. Cognitive and Knowledge Demands, 2. Language and Content Quality, 3. Task Properties, 4. Context, and 5. Ethics, Safety, and Fairness. Applying this framework, we annotate five influential benchmarks -- MMLU, ARC, WinoGrande, HellaSwag, and TruthfulQA -- revealing pronounced internal heterogeneity that is not captured by aggregate accuracy scores. We show how these annotations enable criterion-driven orchestration of composite benchmark subsets across datasets, supporting targeted evaluation of model capabilities such as Reasoning Depth or Ethical Sensitivity. This approach reframes benchmark evaluation as dataset introspection, providing a principled methodology for analyzing and re-composing existing benchmarks to better reflect diverse evaluation needs.

View source

Similar papers

Review Jul 2026

Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications

It is taken that financial LLM validation should be an ongoing system discipline rather than a one-time model scoring exercise, and a research agenda for system-aware benchmarks, agent trace validation, judge alignment protocols, and lifecycle validation standards is concluded.

Burak Payzun, Irem Demirtas, Simona Scala et al. · 0 citations
Book Open access Aug 2026

OmniVul: A Holistic, Multi-Turn Conversational Benchmark for LLM-Based Vulnerability Assessment

With more than 20,000 Common Vulnerabilities and Exposures (CVEs) reported annually, software vulnerabilities represent a critical cybersecurity challenge. This volume has intensified the demand for automated detection and analysis, motivating the integration of large language models (LLMs) for such tasks. However, exi...

Vishnu Teja Kandalam, Viet Duong, Xiaochang Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

This work systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026, using staged screening and automated full-text coding to examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms.

Chao Wang · 0 citations
Conference Aug 2026

Improving the Reliability of LLM Evaluation Metrics via Human-in-the-Loop Validation

Large language models (LLMs) are increasingly evaluated using automated metrics such as ROUGE, BERTScore, and perplexity. However, these scores often fail to reflect real-world usefulness, particularly for tasks requiring complex reasoning or agentic behavior. This paper examines the risks of misaligned LLM evaluation...

Karthik Babu Manam, Vincent Koc, Jamshaid Iqbal Janjua · 0 citations
Conference Open access Sep 2025

DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values

DiverValue-Bench is introduced, a population-aware benchmark for evaluating multi-dimensional value alignment across 74 countries/regions and it is shown that lightweight preference-based fine-tuning with Low-Rank Adaptation and Direct Preference Optimization substantially improves in-domain value alignment while yield...

Yao Liang, Dongcheng Zhao, Fei-Fei Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.