Skip to content

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

Jul 2026 · arXiv.org · Vol abs/2607.09649 · 0 citations · 54 references
Computer Science

TL;DR

ConceptSMILE provides an independent audit layer for evaluating the trustworthiness of concept-based XAI, a model-agnostic perturbation-based auditing framework for evaluating the reliability of concept-based explanations.

Abstract

Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concept-level outputs are not automatically trustworthy. We introduce ConceptSMILE, a model-agnostic perturbation-based auditing framework for evaluating the reliability of concept-based explanations. Rather than replacing SMILE, ConceptSMILE extends its perturbation-based logic from feature- or region-level attribution to the auditing of human-understandable concept explanations. The framework perturbs input regions, measures concept-response shifts, applies locality weighting, and fits an XGBoost surrogate to approximate local concept behaviour. Reliability is assessed through attribution accuracy, surrogate fidelity, faithfulness, stability, and consistency. We evaluate ConceptSMILE on retinal fundus images by comparing MedSAM-derived visual concepts with VLM-based semantic concepts. Results show that reliability varies across concepts and pathways: MedSAM achieves stronger spatial attribution and the highest surrogate fidelity ($R^2 = 0.8503$, $R_w^2 = 0.8465$), while the VLM pathway shows stronger vessel faithfulness and stronger stability under selected artefact conditions. ConceptSMILE provides an independent audit layer for evaluating the trustworthiness of concept-based XAI.

View source

Similar papers

Review Open access Aug 2026

Explainable Artificial Intelligence Technology and Its Applications: Toward Transparent, Human-Centered, and Trustworthy AI

It is demonstrated how XAI is moving from generic explanation visualizations toward domain-sensitive, data-aware, and operationally reliable methods in network security, computer vision, knowledge graphs, and financial decision support.

Xuewen Sun, Lan Tian, Wei-Dong Zhou et al. · 0 citations
Open access Aug 2026

Counterfactual Trust-Aware Explainability for Multimodal AI Decision Systems

The proposed CTAE framework provides a robust and trustworthy explainability solution for high-stakes multimodal AI applications requiring transparent and cognitively reliable decision interpretation.

N. Patil, Afreen Arif · 0 citations
Review Open access Sep 2026

A structured framework for human-AI resemblance: Evidence, inference and perception.

A structured framework in which every resemblance claim specifies the human reference class, the AI system and version, the task and context, the property compared, the measurement relation, the perturbations considered, the uncertainty of the estimate and the inference that the evidence permits is proposed.

Peng Wang, E. Law, Li-Ye Zou et al. · 0 citations
Preprint Aug 2026

Are Concept Bottleneck Models Effective as Decision-Support Systems?

Results show that CBMs, and particularly their interactive component, can improve human-AI team accuracy relative to both unaided human performance and performance with non-interpretable AI support, however, these benefits emerge only under certain conditions: classification tasks perceived as difficult, easily identif...

A. Bogani, Nicola Debole, E. Marconato et al. · 2 citations
Conference Open access Sep 2026

Concept Bottleneck Models for Explainable Decision Making: A Survey of Progress, Taxonomy, and Future Directions

Deep neural networks deliver strong performance but remain opaque, limiting their use in high-stakes domains that require transparency and human oversight. Concept Bottleneck Models (CBMs) address this gap by introducing a human-interpretable concept layer that mediates inputs and decisions, enabling semantic explanati...

Chun-Jiang Wang, Fan Li, Wen-Bo Hu et al. · 2 citations
#artificial intelligence Preprint Aug 2026

Validity-Aware Jailbreak Evaluation for Large Language Models

This work proposes Sequential Epistemic and Action-Level Validation (SEAV), a verification-centric jailbreak evaluation framework that decomposes responses into ordered steps and evaluates both validity and correctness, and shows that enforcing correctness substantially reshapes measured robustness.

Qilong Wu, Sahil Wadhwa, Pranab Mohanty et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.