This work introduces a new approach to steering neural networks based on partial dependence, such that their average response to certain features aligns with specific functional domain knowledge about the problem.
Abstract
Over the last few years, there has been an increased interest in making machine learning models more interpretable. Although a great deal of effort goes into developing techniques for interpreting the interactions learned by a given model, fewer studies focus on assessing the quality of such explanations. Even fewer focus on how to adjust the model to produce explanations faithful to prior knowledge, a process known as explanation-guided learning. Furthermore, most approaches in this area focus on classification problems and usually assume prior knowledge about which input features or regions are most important. In this work, we introduce a new approach to steering neural networks based on partial dependence, such that their average response to certain features aligns with specific functional domain knowledge about the problem. We empirically demonstrate on a range of regression problems, including dynamical systems forecasting, that models whose training has been controlled using our method perform better than unconstrained models and are more data-efficient. Moreover, we highlight that interpretations obtained from the former actually align with the user-provided knowledge, whereas those obtained from the latter do not.
The practically actionable issue that should be addressed when a prediction is wrong is addressed, which is based on the following caveat: Rather than explaining a prediction which is of little-to-no impact given the relatively high likelihood that it is incorrect, the issue of explaining the reasons for the high uncer...
Tameem Adel· International Journal of Art...· 0 citations
A curious phenomenon called mode connectivity, the ability to connect neural networks in the loss surface, defies explanation entirely is elucidates, explains and exploits this special structure in the loss landscape.
This work introduces space explanations, a logic-based notion of explanation that represents sufficient conditions for a neural network to predict a given class over a (potentially large and geometrically complex) subset of the feature space and demonstrates that the interpolation-based explanations are more meaningful...
Faezeh Labbaf, Tomáš Kolárik, Martin Blicha et al.· CI-BD-SOQE@FLoC· 0 citations
This work proposes a novel LAL method for classification that exploits symmetry and independence properties of the active learning problem with an Attentive Conditional Neural Process model and gives the model the ability to adapt to non-standard objectives.
This work proposes a method to sample from models'perceptual prior distributions directly, by steering a generative model to produce stimuli along controllable axes and running Gibbs sampling over that space with the model under study as the judge, and recovers both canonical biases and surprising novel priors invisibl...
Manuel Cherep, Pattie Maes, Nikhil Singh· 2 citations
IB-Forecast is proposed, an inherently interpretable multivariate time-series forecasting framework that decomposes forecasting into a learned periodic component and a residual component computed with explainable masks over input tokens that guarantees high explanation fidelity.