Skip to content

Steering Neural Network Training through Interpretable Constraints Based on Partial Dependence

Jul 2026 · arXiv.org · Vol abs/2607.08641 · 0 citations · 40 references
Computer Science

TL;DR

This work introduces a new approach to steering neural networks based on partial dependence, such that their average response to certain features aligns with specific functional domain knowledge about the problem.

Abstract

Over the last few years, there has been an increased interest in making machine learning models more interpretable. Although a great deal of effort goes into developing techniques for interpreting the interactions learned by a given model, fewer studies focus on assessing the quality of such explanations. Even fewer focus on how to adjust the model to produce explanations faithful to prior knowledge, a process known as explanation-guided learning. Furthermore, most approaches in this area focus on classification problems and usually assume prior knowledge about which input features or regions are most important. In this work, we introduce a new approach to steering neural networks based on partial dependence, such that their average response to certain features aligns with specific functional domain knowledge about the problem. We empirically demonstrate on a range of regression problems, including dynamical systems forecasting, that models whose training has been controlled using our method perform better than unconstrained models and are more data-efficient. Moreover, we highlight that interpretations obtained from the former actually align with the user-provided knowledge, whereas those obtained from the latter do not.

View source

Similar papers

Jul 2026

EXPLAINING UNCERTAINTY ESTIMATES BASED ON GENERATIVE MODELS

The practically actionable issue that should be addressed when a prediction is wrong is addressed, which is based on the following caveat: Rather than explaining a prediction which is of little-to-no impact given the relatively high likelihood that it is incorrect, the issue of explaining the reasons for the high uncer...

Tameem Adel · 0 citations
2026

Using Craig Interpolation for Explanation of Neural Networks (Abstract)

This work introduces space explanations, a logic-based notion of explanation that represents sufficient conditions for a neural network to predict a given class over a (potentially large and geometrically complex) subset of the feature space and demonstrates that the interpolation-based explanations are more meaningful...

Faezeh Labbaf, Tomáš Kolárik, Martin Blicha et al. · 0 citations
Review

Learning Strategies with Attentive Neural Processes

This work proposes a novel LAL method for classification that exploits symmetry and independence properties of the active learning problem with an Attentive Conditional Neural Process model and gives the model the ability to adapt to non-standard objectives.

Tim Bakker, H. van Hoof, Max Welling · 0 citations
#artificial intelligence Preprint Aug 2026

Probing Perceptual Priors of MLLMs via Gibbs Sampling with Interpretable Generative Controls

This work proposes a method to sample from models'perceptual prior distributions directly, by steering a generative model to produce stimuli along controllable axes and running Gibbs sampling over that space with the model under study as the judge, and recovers both canonical biases and surprising novel priors invisibl...

Manuel Cherep, Pattie Maes, Nikhil Singh · 2 citations
Jul 2026

Information Bottleneck Learning for Faithful Time Series Forecasting Explanations

IB-Forecast is proposed, an inherently interpretable multivariate time-series forecasting framework that decomposes forecasting into a learned periodic component and a residual component computed with explainable masks over input tokens that guarantees high explanation fidelity.

Xu Zheng, Wei Cheng, Zhuomin Chen et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.