Skip to content
Preprint

Handling Missingness and Censoring in Dirichlet Models

Jul 2026 · 1 citation
Mathematics

Abstract

Likelihood-based inference for compositional data generally requires fully observed compositions, hindering the direct treatment of missing or censored components on the simplex. In this paper, we develop an expectation-maximisation (EM)-type algorithm for maximum likelihood estimation of the Dirichlet parameters in the presence of missing and censored components under a unified coarsening framework. The Dirichlet distribution---the canonical probability model for compositional data, which plays a role analogous to that of the multivariate normal distribution for unconstrained multivariate data---provides the foundation for our methodology. Our methodology preserves the compositional structure of the data while simultaneously performing parameter estimation and model-based imputation. We evaluate the performance of our estimators and imputations through a simulation study under increasingly complex coarsening mechanisms, including both missing and censored data. We compare our method with an existing model-based approach and a nonparametric alternative. Finally, we illustrate the practical utility of our methodology using mercury speciation data, in which compositions are only partially observed because of detection limits and incomplete speciation. Our results indicate that the Dirichlet distribution provides a suitable model for these data and that our method yields imputations that better preserve the observed compositional structure than competing approaches.

View source

Similar papers

Review Aug 2026

Handling mild outliers and unobserved values in compositional datasets using finite mixtures of mean-parametrised Dirichlet models

Heterogeneous compositional data may be simultaneously affected by missing values and atypical points, posing challenges for both clustering and outlier detection. We develop a mixture model for incomplete compositional data under Huber's contamination model, with contamination defined directly on the simplex and on ob...

J. Pillay, A. Bekker, C. Tortora et al. · 0 citations
Preprint Sep 2026

Model-based estimation and imputation with torus missing values

This paper addresses the problem of parameter estimation and model-based imputation for multivariate circular data lying on a p-dimensional torus in the presence of missing values. Actually, the periodic nature of the sample space invalidates conventional imputation techniques designed for Euclidean data. Then, we prop...

Luca Greco, Lucia Filippozzi, Claudio Agostinelli · 0 citations
Preprint Aug 2026

A New Look at Gaussian Mixtures in the Presence of Missing-at-Random Responses and Covariates

Missing values present a common challenge in statistical modeling, so handling them properly is an important research direction. Among the various mechanisms that can generate missing values, the most common is the missing-at-random (MAR) mechanism, in which the probability of missingness depends only on observed data...

Hung Tong, A. Punzo, C. Tortora · 1 citation · ⚡1
Review Open access Aug 2026

ReX-CoDA: A Compositionally Aware Imputation Framework for Censored Geochemical Data

An essential preprocessing step in geochemical data analysis is the identification and handling of censored data that occur when elemental concentrations fall below analytical detection limits. These censored observations must be estimated in a way that preserves the spatial coherence and multivariate structure of th...

O. Emenaha, H. Basarir, S. Ellefmo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.