Likelihood-based inference for compositional data generally requires fully observed compositions, hindering the direct treatment of missing or censored components on the simplex. In this paper, we develop an expectation-maximisation (EM)-type algorithm for maximum likelihood estimation of the Dirichlet parameters in the presence of missing and censored components under a unified coarsening framework. The Dirichlet distribution---the canonical probability model for compositional data, which plays a role analogous to that of the multivariate normal distribution for unconstrained multivariate data---provides the foundation for our methodology. Our methodology preserves the compositional structure of the data while simultaneously performing parameter estimation and model-based imputation. We evaluate the performance of our estimators and imputations through a simulation study under increasingly complex coarsening mechanisms, including both missing and censored data. We compare our method with an existing model-based approach and a nonparametric alternative. Finally, we illustrate the practical utility of our methodology using mercury speciation data, in which compositions are only partially observed because of detection limits and incomplete speciation. Our results indicate that the Dirichlet distribution provides a suitable model for these data and that our method yields imputations that better preserve the observed compositional structure than competing approaches.
Heterogeneous compositional data may be simultaneously affected by missing values and atypical points, posing challenges for both clustering and outlier detection. We develop a mixture model for incomplete compositional data under Huber's contamination model, with contamination defined directly on the simplex and on ob...
J. Pillay, A. Bekker, C. Tortora et al.· 0 citations
This paper addresses the problem of parameter estimation and model-based imputation for multivariate circular data lying on a p-dimensional torus in the presence of missing values. Actually, the periodic nature of the sample space invalidates conventional imputation techniques designed for Euclidean data. Then, we prop...
Missing values present a common challenge in statistical modeling, so handling them properly is an important research direction. Among the various mechanisms that can generate missing values, the most common is the missing-at-random (MAR) mechanism, in which the probability of missingness depends only on observed data...
Comparisons with MCMC indicate that the variational approximation produces clustering results and parameter estimates that are in close agreement with those obtained by MCMC, while requiring substantially lower computational cost.
An essential preprocessing step in geochemical data analysis is the identification and handling of censored data that occur when elemental concentrations fall below analytical detection limits. These censored observations must be estimated in a way that preserves the spatial coherence and multivariate structure of th...
O. Emenaha, H. Basarir, S. Ellefmo· Mathematical Geosciences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.