Skip to content
Review Open access

A survey of theory-grounded interpretability for deep neural networks

Aug 2026 · Journal of King Saud University: Computer and Information Sciences · Vol 38 · 0 citations · 278 references

Abstract

Deep neural networks achieve state-of-the-art performance across computer vision, natural language processing, and scientific discovery, yet their internal decision-making mechanisms remain difficult to interpret. Post-hoc explanation methods such as SHAP, saliency maps, and attention visualization provide empirical insights but offer limited mathematical guarantees about learning dynamics and representation formation. This survey takes a complementary, theory-grounded perspective, examining interpretability through four classical machine learning (ML) frameworks, each grounded in a mature mathematical theory. These are kernel methods (functional analysis), sparse representations (convex optimization), matrix factorization (linear algebra), and manifold learning (differential geometry). We synthesize and critically evaluate 203 papers (2016–2026), organizing them by the theoretical principle each framework contributes rather than by application domain. Neural tangent kernel (NTK) methods provide rigorous convergence and generalization guarantees but depend on idealized infinite-width assumptions; sparse representation methods yield interpretable-by-design architectures but face reconstruction and approximation trade-offs; matrix factorization reveals an implicit bias toward low-rank structure that helps explain generalization; and manifold learning offers geometric tools for analyzing representation structure, class separability, and generative behavior. Beyond reviewing each framework, we establish explicit mathematical bridges between them, arguing that they offer complementary views of a common phenomenon, namely that gradient-based optimization often drives deep networks toward low-dimensional, interpretable structure. We further situate these frameworks among adjacent interpretability paradigms, including mechanistic, concept-based, and attribution-based methods, and delineate the validity conditions and failure modes that determine when classical theory yields reliable rather than misleading explanations. The review concludes with open problems and directions for scalable, theory-grounded interpretability.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.