Aug 2026· 2026 7th International Conference on Big Data Analytics and Practices (IBDAP)· pp. 1-6· 0 citations· 19 references
Abstract
Large language models (LLMs) are increasingly evaluated using automated metrics such as ROUGE, BERTScore, and perplexity. However, these scores often fail to reflect real-world usefulness, particularly for tasks requiring complex reasoning or agentic behavior. This paper examines the risks of misaligned LLM evaluation and identifies key failure modes where metrics reward outputs that are semantically incorrect or provide little value to users. We propose a multi-layer evaluation framework that combines statistical metrics, model-based judges (for example, G-Eval and SelfCheckGPT), and human-in-the-loop (HITL) review. Using stress-tested datasets and perturbation analyses, we expose vulnerabilities in widely used metrics and quantify their false-alignment rates. To mitigate these issues, we introduce a disagreement-driven evaluation loop and a multi-metric aggregation module that prioritizes safety and factuality. Overall, our framework offers practical design principles for building more trustworthy LLM evaluation pipelines and for advancing task-aligned, semantics-focused measurement standards.
A lightweight quality-assessment protocol is presented for LLM-generated synthetic training data and applied to 13,579 synthetic user reviews generated from GitHub issues across four open-source Android applications, high-lighting the need for hybrid human-AI verification when synthetic data is used in security-critica...
Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without...
Mario Sanz-Guerrero, Katharina von der Wense· 1 citation
This work systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026, using staged screening and automated full-text coding to examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms.
This work proposes an explanation-aware safety framework that augments binary harmfulness detection with structured, human-interpretable explanations capturing severity, strategies, trigger spans, ratio-nales, and derived safety factors, and introduces a human–LLM hybrid annotation and canonicaliza-tion pipeline.
Sunghee Dong, Sungwon Yi, K. Bae et al.· Proceedings of the Thirty-Fi...· 0 citations
Prompt sensitivity is widely treated as a model robustness deficiency. Yet the extent to which measured sensitivity reflects genuine model instability, rather than artifacts of the evaluation method used to measure it, remains largely underexplored. We introduce Evaluation-Attributable Sensitivity (EAS), a per-instance...
Sayumi Muthukumarana, Buddhi Wijenayake, R. Godaliyadda et al.· Moratuwa Engineering Researc...· 0 citations
This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...
Himil Vasava, Ming-Zhou Jiang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.