Skip to content
Open access

A framework for rigorous evaluation of model-agnostic explainability methods: multi-metric statistical benchmarking, operational protocol, and reproducibility

Sep 2026 · Revista de investigación multidisiplinaria, Iberoamericana · 0 citations · 36 references

TL;DR

A modular benchmarking framework centered on quantitative XAI quality metrics: fidelity, stability, sparsity, computational cost, and faithfulness gap, plus an explicit method for operating the framework end-to-end is presented, plus an explicit method for operating the framework end-to-end.

Abstract

Evaluating explainability methods requires more than a single faithfulness proxy. We present a modular benchmarking framework centered on quantitative XAI quality metrics: fidelity, stability, sparsity, computational cost, and faithfulness gap, plus an explicit method for operating the framework end-to-end. On the UCI Adult benchmark, we use a staged evidence protocol: a calibration/reproducibility stage (EXP1) followed by a primary com-parative/robustness benchmark (EXP2); the current merged recovery snapshot contains 299 committed result artifacts (99.7% artifact coverage) plus a 30-row SHAP recovery batch, yielding 275 analyzable unique runs out of 300 planned cells (91.7%). Across complete model-size blocks (5 models, N ∈ {50, 100, 200}), Friedman tests indicate significant method differences for fidelity (χ2 = 42.12, p = 3.78 × 10−9), stability (χ2 = 40.68, p = 7.65 × 10−9), sparsity (χ2 = 35.64, p = 8.92 × 10−8), faithfulness gap (χ2 = 45.00, p = 9.25 × 10−10), and runtime (χ2 = 30.44, p = 1.12 × 10−6). SHAP leads on fidelity/stability, DiCE leads on sparsity, and LIME remains fastest overall. We release the framework, operation protocol, and artifacts with explicit data-quality caveats for reproducible benchmark use under a quantitative-only claim scope.

Read PDF

Similar papers

#machine learning Preprint Oct 2026

Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability

Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circ...

Chu-Qin Geng, Li Zhang, Hao-Lin Ye et al. · 0 citations
Preprint Aug 2026

A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification

An auditing protocol is constructed that measures two properties of any post-hoc explainer: robustness (how stable the explanation is under input perturbation) and fidelity (whether the features deemed important actually drive the model's prediction).

Rosa Elysabeth Ralinirina, J. Ralaivao, Niaiko Michaël Ralaivao et al. · 0 citations
Open access Aug 2026

condfair: An R Package for Ability-Conditioned Fairness and Explanation Diagnostics in Automated Scoring.

Fairness in automated scoring is typically evaluated with a single global statistic contrasting a focal and reference group-an approach that can either mask a disparity that changes sign across the ability range, or overstate one by conflating it with genuine ability differences between groups (impact). We introduce co...

Tri Zahra Ningsih, Aman Aman, Ahmad Nasrulloh · 1 citation
Preprint Aug 2026

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.

V. Rodionov, Shamil Assylbekov · 0 citations
#artificial intelligence Preprint Sep 2026

SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores

Scientific coding agents produce interdependent code, results, figures, and claims, yet evaluating final outputs alone does not establish whether their conclusions are scientifically supported. We formulate evidence-grounded multimodal scientific analysis, requiring agents to produce executable analyses and claims supp...

Bowen Liu, Shuo Nie, Bo-Dong Du et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.