Skip to content

LiD-GLM: Lipschitz-constrained Deep Generalized Linear Models

Aug 2026 · 0 citations · 33 references
Mathematics Computer Science

TL;DR

The method proposed leverages invertible residual neural networks (i-ResNets) to equip generalized linear models with both nonlinear parameter estimation and a flexible correction of their distributional assumptions while always retaining stochastic monotonicity of the modeled distribution in the (formerly linear) predictor.

Abstract

The combination of traditional statistical models and neural network (NN) components into semi-structured hybrid models is an intriguing approach to construct models that, ideally, combine traditional interpretability with the unprecedented flexibility of NNs. In order to preserve interpretability, it is usually necessary to restrict the NN components to prevent them from dominating the model. However, existing methods that enforce structural constraints on their NN components severely limit their models'flexibility; in contrast, methods that only enforce weak, indirect constraints lose meaningful interpretability. The method we propose therefore leverages invertible residual neural networks (i-ResNets) to equip generalized linear models with both nonlinear parameter estimation and a flexible correction of their distributional assumptions while always retaining stochastic monotonicity of the modeled distribution in the (formerly linear) predictor. The i-ResNets correspond to a controlled deviation from identity and by constraining their Lipschitz constant one can rigorously limit and quantify how far the hybrid model deviates from its traditional counterpart. This enables a user-specifiable compromise between flexibility and interpretability without limiting the structure of nonlinear and interaction effects that can be learned. Furthermore, we develop specific inherent interpretation techniques for our model and enforce model identifiability through an adapted post-hoc orthogonalization.

View source

Similar papers

Review Open access Sep 2026

An Introduction to Stochastic Deep Learning

Deep neural networks (DNNs) have achieved remarkable success in prediction, but their deterministic formulation makes many statistical inference tasks difficult. StoNet, short for stochastic neural network, addresses this limitation by reformulating a DNN as a probabilistic latent‐variable model, in which the outputs o...

Fa-Ming Liang · 0 citations
#artificial intelligence Preprint Sep 2026

A unified framework for global and local interpretability using adaptive derivative-ordered random explanation

The interpretability of complex machine learning models is of paramount importance, especially in real-world high-stakes domains such as healthcare and finance. However, existing post-hoc interpretability methods suffer from inherent limitations: fragmented analytical processes, inadequate capacity to model nonlinear f...

Le-Men Chao, Ming Lei, An-Ran Fang · 0 citations
Preprint Sep 2026

From Good Starts to Optimal Inference: Generalized Latent Factor Models with Missingness and Implicit Regularization

A theory is developed that connects a computationally tractable nonconvex procedure directly to statistical inference for nonlinear latent factor models with exponential-family links and partially observed entries with severe missingness, weak low-rank signals, and diminishing local curvature.

Cheng-Zhu Huang, Yu-Qi Gu · 0 citations
Book Open access Aug 2026

Sparse Additive Models for Domain Generalization

This work incorporates an additive structure into the DG framework and employs ℓq,1 -norm regularization to induce sparsity, thereby enabling structured feature selection and enhancing interpretability and presents two distinct realizations: an additive kernel-based formulation and a neural additive model-based approac...

Jia-Yi Wang, Han Li · 0 citations
#machine learning Preprint Sep 2026

A Distributional Optimisation Perspective on Combining Models in Deep Learning

Combining predictions from different models can improve performance at machine learning tasks, but the training of the individual models and the rule used to combine them are typically chosen separately, and by ad hoc means. Recent advances in distributional optimisation (i.e. where the optimisation occurs over the set...

Cong-Ye Wang, Yan-Kai Lin, Zhe-Yang Shen et al. · 0 citations
Preprint Aug 2026

A Generalized Ridge Regression and Convolutional LASSO

It is proved that the extracted trend is invariant across representations, and a closed-form Bregman-type divergence quantifying their disagreement under a shared coefficient vector is obtained, and this divergence vanishes as $\lambda\to\infty$.

Shintaro Yoshizawa · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.