Skip to content

All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language Models

Jul 2026 · arXiv.org · Vol abs/2607.09502 · 0 citations
Computer Science

TL;DR

The Rashomon Explanation paradigm is introduced, which builds a set of faithful, prediction-guiding explanations rather than a single one, and it is proved that this set is generally non-empty and that explanation fidelity bounds the performance of the models it guides.

Abstract

Explaining machine-learning models is increasingly important for decision-making and consumer trust, yet it is widely believed to come at a cost: existing Explainable AI (XAI) methods suffer from a persistent accuracy-explainability trade-off. We argue that this trade-off is not fundamental, but an artifact of treating explanation and prediction as separate objectives; when properly coupled, they become complementary, so that equipping a model to explain itself improves, rather than degrades, its accuracy. We introduce the Rashomon Explanation paradigm, which builds a set of faithful, prediction-guiding explanations rather than a single one, and prove that this set is generally non-empty and that explanation fidelity bounds the performance of the models it guides. To explore this set, we propose RashomonLLM, an Explanation-Prediction-Reflection agentic workflow that generates explanations in natural language by iteratively aligning them with predictions, and we prove it converges and recovers the full set. Across customer-churn classification, clinical survival regression, and industrial click-through prediction on large-scale live-streaming logs, RashomonLLM significantly outperforms state-of-the-art prediction and XAI baselines on both accuracy and explanation quality, with gains driven by explanation fidelity and robust to distribution shifts, temporal splits, and seeds. Our framework thus advances business performance while laying the groundwork for consumer trust.

View source

Similar papers

Review Jul 2026

ExplainBench: Evaluating Code Explanations from Agents

This work proposes ExplainBench, a benchmark to automatically evaluate explanations from coding agents, based on the intuition that informative explanations should enable an LLM to correctly answer questions, allowing quantitative comparison of explanation quality between agents.

Zhiyuan Pan, Sung-Min Kang, Imam Nur Bani Yusuf et al. · 0 citations
#machine learning Preprint Sep 2026

Evaluating Explanation Methods by the Predictors They Induce

Explanations of machine learning models are usually judged by criteria that are hard to compare. We propose a simpler test: if an explanation really describes how a model uses its features, it should be possible to rebuild the model's predictions from it. We turn each explanation into a predictor by reading each featur...

Jacob Selbæk, Hugo L. Hammer · 0 citations
#artificial intelligence Book Open access Aug 2026

The Utility of LLMs in Recommender Systems Explanation Evaluation

Explanations play a crucial role in creating trustworthy recommender systems (RS), yet choosing a good explanation method presents challenges. Many explanation methods exist, but little guidance exists on which is best for which setting. Existing explanation generation methods often produce abstract outputs that requir...

Kathrin Wardatzky, Oana Inel, Luca Rossetto et al. · 0 citations
Preprint Aug 2026

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

The Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable.

Zi-Yue Wang, Aomufei Yuan, Yi-Ran Yao et al. · 0 citations
#machine learning Preprint Sep 2026

Do Reasoning Representations Help Humans Evaluate LLM Outputs?

This work conducts a controlled human study of six reasoning formats across tasks of varying complexity, supported by a web-based framework that randomizes task domains, problem instances, and representation order.

Jaewoo Lim, Sungbok Shin, San Hong · 0 citations
Preprint Aug 2026

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.

V. Rodionov, Shamil Assylbekov · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.