Skip to content
Conference

Auditing Industrial Wind RAG Evaluation: Cross-Family Scoring and Fair Question Selection Reveal Graph-Enhanced Retrieval Advantages

Aug 2026 · 2026 8th International Conference on System Reliability and Safety Engineering (SRSE) · pp. 854-859 · 0 citations · 21 references

Abstract

Retrieval-augmented generation (RAG) benchmarks increasingly use a single large language model (LLM) judge, whose bias can hide differences between retrieval methods. We audit a wind-turbine operation-and-maintenance benchmark with 173 questions, eight methods, and LLM-generated reference answers. Reviewing 65 items identifies five recurring scoring defects and a 28% (17/60) mis-kill rate for a graph-enhanced fusion method. We repair the evaluation with a four-dimensional, question-type-routed protocol and a mechanical same-source subset rule. On 47 clean questions, the fusion method scores 0.805 versus BM25 at 0.721 (paired Wilcoxon $p=0.0019)$ and ranks first for all three question types. A protocol-by-subset control attributes the reversal to scoring rather than question selection, and a second model family reproduces the method ranking (Spearman $\rho=0.952)$. Limited ablation, weight sensitivity, and API telemetry characterize the contribution, stability, and cost of the repair. Because reference answers and most diagnostic judgments remain LLM-based, we report a protocol-conditional advantage rather than an absolute ranking.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.