Skip to content
Open access

Bridging trust and performance in intelligent systems: Hybrid explainable AI approaches for interpreting large language models

Sep 2026 · PLoS ONE · Vol 21, pp. e0343472 · 0 citations · 52 references
Medicine

TL;DR

A hybrid explainability framework that integrates saliency-based attribution, causal reasoning, and user-centered visualization into a unified, efficiency-aware pipeline is introduced, positioning hybrid XAI as a pathway toward responsible LLM adoption.

Abstract

Large Language Models (LLMs) achieve state-of-the-art performance across natural language processing tasks but remain opaque, limiting adoption in high-stakes domains that demand accountability and transparency. This paper introduces a hybrid explainability framework that integrates saliency-based attribution, causal reasoning, and user-centered visualization into a unified, efficiency-aware pipeline. Unlike prior single-method approaches such as LIME, SHAP, or attention visualization, the framework provides explanations that are both technically faithful and accessible to human evaluators. The framework was systematically evaluated on benchmark datasets (GLUE, SQuAD, IMDB, and domain-specific corpora) and tested on representative architectures (BERT, T5, GPT, and LLaMA). Results show up to a 15–20% improvement in fidelity compared to attention-based methods. Fidelity was measured using standardized insertion and deletion metrics across all benchmark datasets using a consistent evaluation protocol, ensuring objective and comparable assessment of explanation faithfulness across different LLM architectures. The proposed framework also achieved higher clarity and trust ratings in user studies while introducing less than 25% additional computational overhead. Case studies in sentiment analysis and question answering further demonstrate that hybrid explanations produce precise, intuitive reasoning paths that outperform existing baselines. The main contributions are: (1) a multi-method pipeline that reconciles the trade-off between faithfulness and interpretability; (2) a human-centered evaluation showing hybrid explanations are more trustworthy than single techniques; and (3) an efficiency-aware design indicating the potential suitability of the proposed framework for practical applications in domains such as healthcare, finance, and law. By aligning methodological rigor with societal and regulatory demands, this study advances both the practice and theory of explainable AI, positioning hybrid XAI as a pathway toward responsible LLM adoption.

Read PDF

Similar papers

Book Open access Aug 2026

Interpretability in the Era of Large Language Models: Mechanistic Methodology, Empirical Practices, and Applications

This tutorial provides a comprehensive, end-to-end view of LLM interpretability, transitioning from microscopic neural analysis to macroscopic application and deployment, and explores how these interpretability paradigms scale and inspire the design of frontier architectures, agentic systems, and thinking models.

Wei Zhang, Zheng-Fu He, Lu-Lu Zhang et al. · 0 citations
Open access Aug 2026

Towards Trustworthy Large Language Models

An integrated conceptual frame-work that couples attention- and perturbation-based explainability with lightweight hallucination-detection signals and token-efficient inference strategies is presented, and a set of cross-cutting consistency metrics are instrumented with a set of cross-cutting consistency metrics.

Sakshi Parate, Shreyans Sanyal · 0 citations
Preprint Sep 2026

Transferring Visual Explanations: How Cross-Architecture Knowledge Distillation Affects Model Interpretability

Deploying efficient neural networks is essential in resource-constrained environments, yet compact models often sacrifice interpretability - a critical in safety-critical domains such as autonomous driving and medicine. This study investigates whether Knowledge Distillation transfers the spatial feature attribution of...

Aleks Czufarow, I. Babin · 0 citations
#machine learning Preprint Sep 2026

Do Reasoning Representations Help Humans Evaluate LLM Outputs?

This work conducts a controlled human study of six reasoning formats across tasks of varying complexity, supported by a web-based framework that randomizes task domains, problem instances, and representation order.

Jaewoo Lim, Sungbok Shin, San Hong · 0 citations
#machine learning Preprint Sep 2026

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...

Himil Vasava, Ming-Zhou Jiang · 0 citations
Conference Open access Sep 2026

Explaining Jailbreaks: Structured and Interpretable Safety Assessment for Large Language Models

This work proposes an explanation-aware safety framework that augments binary harmfulness detection with structured, human-interpretable explanations capturing severity, strategies, trigger spans, ratio-nales, and derived safety factors, and introduces a human–LLM hybrid annotation and canonicaliza-tion pipeline.

Sunghee Dong, Sungwon Yi, K. Bae et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.