Skip to content
Conference

Improving the Reliability of LLM Evaluation Metrics via Human-in-the-Loop Validation

Aug 2026 · 2026 7th International Conference on Big Data Analytics and Practices (IBDAP) · pp. 1-6 · 0 citations · 19 references

Abstract

Large language models (LLMs) are increasingly evaluated using automated metrics such as ROUGE, BERTScore, and perplexity. However, these scores often fail to reflect real-world usefulness, particularly for tasks requiring complex reasoning or agentic behavior. This paper examines the risks of misaligned LLM evaluation and identifies key failure modes where metrics reward outputs that are semantically incorrect or provide little value to users. We propose a multi-layer evaluation framework that combines statistical metrics, model-based judges (for example, G-Eval and SelfCheckGPT), and human-in-the-loop (HITL) review. Using stress-tested datasets and perturbation analyses, we expose vulnerabilities in widely used metrics and quantify their false-alignment rates. To mitigate these issues, we introduce a disagreement-driven evaluation loop and a multi-metric aggregation module that prioritizes safety and factuality. Overall, our framework offers practical design principles for building more trustworthy LLM evaluation pipelines and for advancing task-aligned, semantics-focused measurement standards.

View source

Similar papers

Review

Judging the LLM Judges: A Human-Centric Validation of LLM-Generated Training Data for Software Retrieval

A lightweight quality-assessment protocol is presented for LLM-generated synthetic training data and applied to 13,579 synthetic user reviews generated from GitHub issues across four open-source Android applications, high-lighting the need for hybrid human-AI verification when synthetic data is used in security-critica...

Ogtay Hasanov, Saad Ezzini · 0 citations
#natural language process... Preprint Sep 2026

Calibration as a First-Class Criterion in LLM Evaluation

Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without...

Mario Sanz-Guerrero, Katharina von der Wense · 1 citation
#artificial intelligence Preprint Sep 2026

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

This work systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026, using staged screening and automated full-text coding to examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms.

Chao Wang · 0 citations
Conference Open access Sep 2026

Explaining Jailbreaks: Structured and Interpretable Safety Assessment for Large Language Models

This work proposes an explanation-aware safety framework that augments binary harmfulness detection with structured, human-interpretable explanations capturing severity, strategies, trigger spans, ratio-nales, and derived safety factors, and introduces a human–LLM hybrid annotation and canonicaliza-tion pipeline.

Sunghee Dong, Sungwon Yi, K. Bae et al. · 0 citations
Conference Aug 2026

Prompt Sensitivity or Evaluation Artifact? A Task-Aware Analysis for Large Language Models

Prompt sensitivity is widely treated as a model robustness deficiency. Yet the extent to which measured sensitivity reflects genuine model instability, rather than artifacts of the evaluation method used to measure it, remains largely underexplored. We introduce Evaluation-Attributable Sensitivity (EAS), a per-instance...

Sayumi Muthukumarana, Buddhi Wijenayake, R. Godaliyadda et al. · 0 citations
#machine learning Preprint Sep 2026

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...

Himil Vasava, Ming-Zhou Jiang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.