Skip to content

Distributionally Faithful Imputation via Positive Semi-Definite Kernel Density Estimation

Jul 2026 · arXiv.org · Vol abs/2607.07767 · 0 citations · 37 references
Mathematics Computer Science

TL;DR

This work recast imputation under missing completely at random (MCAR) as density estimation from masked observations: estimate a distribution whose observed marginals exactly match those in the data.

Abstract

Missing values undermine statistical inference and machine learning pipelines, yet most imputation methods rely on heuristics or restrictive parametric assumptions that ignore the joint data distribution. We recast imputation under missing completely at random (MCAR) as density estimation from masked observations: estimate a distribution whose observed marginals exactly match those in the data. Leveraging positive semi definite (PSD) kernel densities we obtain a convex empirical risk problem with closed form marginals, solvable by a Newton interior point method. The resulting PSD Impute model yields both single and multiple imputations from the same fitted density, enjoys statistical consistency with fast adaptive excess risk beating the curse of dimensionality for very regular probabilities. Preliminary experiments on one synthetic and eleven real world datasets already indicate competitive distributional accuracy compared with popular imputation baselines, suggesting strong practical promise.

View source

Similar papers

Preprint Aug 2026

Asymptotics of Nonparametric Estimation under General Non-monotone MAR Missingness: A Nonparametric Maximum Likelihood Approach

Missing data constitute a pervasive challenge in empirical research. Consequently, there is an ever-growing number of methods designed to address this challenge, with multiple imputation and inverse probability weighting the dominant strategies. Despite this, theoretical guarantees remain limited, particularly in the c...

Yating Zou, Hui-Min Hu, Jeffrey Näf · 0 citations
Preprint Jul 2026

Handling Missingness and Censoring in Dirichlet Mixture Models

Incomplete compositional data analysis faces a fundamental limitation: likelihood-based methods for compositional models generally require fully observed compositions, making it difficult to accommodate missing or censored proportions directly on the simplex. Consequently, analysts often discard partially observed comp...

J. Pillay, A. Bekker, C. Tortora et al. · 0 citations
Preprint Sep 2026

From Good Starts to Optimal Inference: Generalized Latent Factor Models with Missingness and Implicit Regularization

A theory is developed that connects a computationally tractable nonconvex procedure directly to statistical inference for nonlinear latent factor models with exponential-family links and partially observed entries with severe missingness, weak low-rank signals, and diminishing local curvature.

Cheng-Zhu Huang, Yu-Qi Gu · 0 citations
Preprint Jul 2026

From dense grids to valid inference: Accounting for regularization bias in nonparametric random coefficient models

This paper develops an inference procedure for average functionals of random-coefficient distributions, such as mean willingness-to-pay and average elasticities, when the distribution is estimated nonparametrically using the penalized fixed-grid estimator of Heiss, Hetzenecker, and Osterhaus (2022). We establish asympt...

Ling-Yan Kong, M. Osterhaus, Michael Pen · 0 citations
Preprint Aug 2026

Handling Missing Data in Probabilistic Regression Trees

Probabilistic Regression Trees (PRTrees) are a smooth and consistent alternative to classical regression trees, producing continuous predictions through probabilistic split assignments. This paper extends the PRTree framework to accommodate missing predictor values directly during tree construction, eliminating the nee...

T. S. Prass, A. Neimaier, G. Pumi · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.