Skip to content
Preprint

A New Look at Gaussian Mixtures in the Presence of Missing-at-Random Responses and Covariates

Aug 2026 · 1 citation · ⚡ 1 influential · 46 references
Mathematics

Abstract

Missing values present a common challenge in statistical modeling, so handling them properly is an important research direction. Among the various mechanisms that can generate missing values, the most common is the missing-at-random (MAR) mechanism, in which the probability of missingness depends only on observed data and not on unobserved data. This paper addresses the problem of estimating a multivariate linear regression model with multiple random covariates in the presence of MAR values in both the response and covariate spaces using a maximum likelihood (ML) framework. The proposed methodology models the joint distribution of responses and covariates through a conditional-marginal factorization of a multivariate Gaussian distribution. This formulation can be interpreted as a reparameterization of the multivariate normal distribution when the variables can be naturally partitioned into responses and covariates. Parameter estimation is performed using the expectation-maximization (EM) algorithm, which facilitates the imputation of missing values while preserving the distinct roles of responses and covariates. We extend this framework to the model-based clustering setting by considering a mixture of multivariate linear regressions with multiple random covariates. This extension enables soft clustering under incomplete data and accommodates MAR values in both the multivariate responses and covariates. Hence, it represents one of the most general model-based clustering solutions for regression data currently available in the literature. The effectiveness of the methodology is demonstrated through a simulation study, and the advantages of the proposed reparameterization are illustrated using the Automobile dataset, which contains missing values.

View source

Similar papers

Preprint Aug 2026

A Unified Framework for Heterogeneity, Contamination, and Missing Data in Multivariate Regression

Missing values, atypical observations, and heterogeneity across latent groups are common sources of complexity in regression data. The contaminated Gaussian cluster-weighted model (CG-CWM) provides a natural framework for handling atypical observations, including outliers and leverage points, in model-based clustering....

Hung Tong, C. Tortora, A. Punzo · 0 citations
Preprint Aug 2026

Asymptotics of Nonparametric Estimation under General Non-monotone MAR Missingness: A Nonparametric Maximum Likelihood Approach

Missing data constitute a pervasive challenge in empirical research. Consequently, there is an ever-growing number of methods designed to address this challenge, with multiple imputation and inverse probability weighting the dominant strategies. Despite this, theoretical guarantees remain limited, particularly in the c...

Yating Zou, Hui-Min Hu, Jeffrey Näf · 0 citations
Review Aug 2026

Handling mild outliers and unobserved values in compositional datasets using finite mixtures of mean-parametrised Dirichlet models

Heterogeneous compositional data may be simultaneously affected by missing values and atypical points, posing challenges for both clustering and outlier detection. We develop a mixture model for incomplete compositional data under Huber's contamination model, with contamination defined directly on the simplex and on ob...

J. Pillay, A. Bekker, C. Tortora et al. · 0 citations
Review Open access Aug 2026

Bayesian model comparison for random effects probit models with missing covariates

For model comparison in random effects probit models with incompletely observed covariates, this paper develops a Bayesian data-augmentation workflow in which latent Gaussian responses, random effects, and missing covariate values are updated within a common augmented sampling scheme. Because specifying a fully param...

Michael Bergrab · 0 citations
Preprint Sep 2026

Quasi-randomization-based inference for multivariate Mann-Whitney effects under random missingness

Marginal Mann-Whitney effects are widely used across various fields of research, and extensions of this estimand have been developed in many directions in statistical methodology. In this paper, we focus on an extensions for repeated measurements and factorial designs subject to randomly missing data. In a previous wor...

Dennis Dobler, Jörg-Tobias Kuhn, L. Amro et al. · 0 citations
Preprint Sep 2026

Model-based estimation and imputation with torus missing values

This paper addresses the problem of parameter estimation and model-based imputation for multivariate circular data lying on a p-dimensional torus in the presence of missing values. Actually, the periodic nature of the sample space invalidates conventional imputation techniques designed for Euclidean data. Then, we prop...

Luca Greco, Lucia Filippozzi, Claudio Agostinelli · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.