A framework for rigorous evaluation of model-agnostic explainability methods: multi-metric statistical benchmarking, operational protocol, and reproducibility
Sep 2026· Revista de investigación multidisiplinaria, Iberoamericana· 0 citations· 36 references
TL;DR
A modular benchmarking framework centered on quantitative XAI quality metrics: fidelity, stability, sparsity, computational cost, and faithfulness gap, plus an explicit method for operating the framework end-to-end is presented, plus an explicit method for operating the framework end-to-end.
Abstract
Evaluating explainability methods requires more than a single faithfulness proxy. We present a modular benchmarking framework centered on quantitative XAI quality metrics: fidelity, stability, sparsity, computational cost, and faithfulness gap, plus an explicit method for operating the framework end-to-end. On the UCI Adult benchmark, we use a staged evidence protocol: a calibration/reproducibility stage (EXP1) followed by a primary com-parative/robustness benchmark (EXP2); the current merged recovery snapshot contains 299 committed result artifacts (99.7% artifact coverage) plus a 30-row SHAP recovery batch, yielding 275 analyzable unique runs out of 300 planned cells (91.7%). Across complete model-size blocks (5 models, N ∈ {50, 100, 200}), Friedman tests indicate significant method differences for fidelity (χ2 = 42.12, p = 3.78 × 10−9), stability (χ2 = 40.68, p = 7.65 × 10−9), sparsity (χ2 = 35.64, p = 8.92 × 10−8), faithfulness gap (χ2 = 45.00, p = 9.25 × 10−10), and runtime (χ2 = 30.44, p = 1.12 × 10−6). SHAP leads on fidelity/stability, DiCE leads on sparsity, and LIME remains fastest overall. We release the framework, operation protocol, and artifacts with explicit data-quality caveats for reproducible benchmark use under a quantitative-only claim scope.
Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circ...
Chu-Qin Geng, Li Zhang, Hao-Lin Ye et al.· 0 citations
The results suggest that benchmark scores should be interpreted together with their scaffolding level, scoring criterion, and reliability profile, providing a practical framework for more valid evaluation of LLM agents.
Yong-Hong Zhang, Shadi Motaali, Vu Phong Dinh et al.· 2 citations
An auditing protocol is constructed that measures two properties of any post-hoc explainer: robustness (how stable the explanation is under input perturbation) and fidelity (whether the features deemed important actually drive the model's prediction).
Rosa Elysabeth Ralinirina, J. Ralaivao, Niaiko Michaël Ralaivao et al.· 0 citations
Fairness in automated scoring is typically evaluated with a single global statistic contrasting a focal and reference group-an approach that can either mask a disparity that changes sign across the ability range, or overstate one by conflating it with genuine ability differences between groups (impact). We introduce co...
Tri Zahra Ningsih, Aman Aman, Ahmad Nasrulloh· Applied Psychological Measur...· 1 citation
A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.
Scientific coding agents produce interdependent code, results, figures, and claims, yet evaluating final outputs alone does not establish whether their conclusions are scientifically supported. We formulate evidence-grounded multimodal scientific analysis, requiring agents to produce executable analyses and claims supp...
Bowen Liu, Shuo Nie, Bo-Dong Du et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.