Skip to content
Review Open access

Cross-Domain Faithfulness Evaluation of SHAP and Attention-Based Explanations in Transformer NLP Models

Jul 2026 · Journal of Computing Theories and Applications · Vol 4, pp. 146-163 · 1 citation · 31 references

TL;DR

The results demonstrate that superior predictive performance does not necessarily correspond to higher explanation faithfulness or stronger cross-domain stability, and highlight the importance of jointly evaluating predictive performance, explanation faithfulness, and explanation robustness when developing trustworthy transformer-based NLP systems.

Abstract

Transformer-based models such as BERT, RoBERTa, DistilBERT, and DeBERTa have achieved remarkable performance across a wide range of natural language processing (NLP) tasks. However, their decision-making processes remain difficult to interpret, particularly in high-risk applications such as hate speech detection, where unreliable explanations may undermine model transparency, trust, and accountability. This study investigates whether explainability methods remain faithful and stable under domain shift in transformer-based text classification. Four transformer architectures were fine-tuned and evaluated on two linguistically distinct datasets: IMDb Movie Reviews and Hate Speech Offensive. Model performance and explanation quality were assessed using classification accuracy, macro F1-score, top-k token-removal faithfulness analysis, and cross-domain Spearman rank correlation. Experimental results show that DeBERTa achieved the highest classification performance, reaching accuracies of 95.6% on IMDb and 91.3% on Hate Speech. Across all evaluated models and datasets, SHAP consistently produced higher faithfulness scores than attention-based explanations. Cross-domain analysis further revealed reduced agreement between SHAP and attention-based explanations under domain shift, indicating lower explanation consistency across linguistically distinct domains. Qualitative error analysis further showed that implicit sentiment, sarcasm, and domain-specific slang remain major sources of prediction errors. Overall, the results demonstrate that superior predictive performance does not necessarily correspond to higher explanation faithfulness or stronger cross-domain stability. These findings highlight the importance of jointly evaluating predictive performance, explanation faithfulness, and explanation robustness when developing trustworthy transformer-based NLP systems.

Read PDF

Similar papers

Open access 2026

An Interpretable Hybrid Ensemble Model for Paraphrase Detection With SHAP-Based Explanations

Paraphrase detection is a fundamental task in natural language processing, typically addressed using high-performing deep neural models that lack interpretability. Although transformer-based and hybrid systems attain high precision, their decision-making procedures are still opaque, which limits their dependability in practical applications. The explainable hybrid framework for paraphrase detection presented in this study combines transformer-based models, semantic similarity models, and recurrent neural networks into a single voting architecture. Token-level attribution of model predictions is provided by SHapley Additive exPlanations (SHAP), which are integrated to improve transparency. Furthermore, a semantic-aware optimization method is used to steer model behavior in the direction of significant similarity patterns. The proposed method is evaluated on the QQP and MRPC, and PAWS benchmark datasets, achieving accuracies of 85.4%, 88.24%, and 94.36% on the QQP, MRPC, and PAWS benchmark datasets respectively. demonstrating strong performance and generalization capability. The results show that the framework provides a balanced trade-off between accuracy and interpretability for paraphrase detection systems.

Dalia Sameh, Ahmed El-Sawy, Eslam Amer et al. · 0 citations
Open access Aug 2026

Assessing reliability of BERT-based models on question answering tasks

This study evaluates the reliability of four BERT-based models by assessing response stability under two conditions: internal model variations induced via Monte Carlo Dropout (MCD) and input perturbations through paraphrasing.

Pooja Yadav, Priyanka Harjule, Basant Agarwal et al. · 0 citations
Open access 2026

Advancing Machine-generated Text Detection: A Comprehensive Evaluation of Transformer-based Models

Test set results show that Decoding-Enhanced Bert with Disentangled Attention (DeBERTa) achieves the highest macro F1 − Score of 85.48%, surpassing the previously top-ranked Multi-Task Learning (MTL) system, which attains a macro F1 of 83.07%.

Batyr Sharimbayev, S. Kadyrov · 0 citations
#machine learning Preprint Sep 2026

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and explicit token-level modification maps, and a four-experiment battery of causal tracing.

Himil Vasava, Mingzhou Jiang · 0 citations
Open access Aug 2026

CVI-Validated Indo-Transformer Framework for Intelligent Cooperative Supervision

It is demonstrated that CVI-validated ensemble GenAI can construct consistent labels for low-resource administrative texts and that IndoBERT provides the strongest and most stable generalization for cooperative supervision classification.

Syahroni Hidayat, Afriani Fajar Navissaturrisqi, Feddy Setio Pribadi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.