ReX-CoDA: A Compositionally Aware Imputation Framework for Censored Geochemical Data
Abstract
An essential preprocessing step in geochemical data analysis is the identification and handling of censored data that occur when elemental concentrations fall below analytical detection limits. These censored observations must be estimated in a way that preserves the spatial coherence and multivariate structure of the dataset. Conventional imputation approaches, such as constant-value substitution or distribution-specific parametric models like the multiplicative lognormal method, often introduce biased estimates, distort variance structure, and weaken inter-element correlations. Modern machine-learning (ML) models, on the other hand, can capture the nonlinear dependencies typical of geochemical systems; however, they are not inherently designed to accommodate the compositional constraints associated with censored observations. This paper therefore introduces ReX-CoDA (Regression-based Flexible-Limit Compositional Data Imputation) as a model-agnostic iterative framework that integrates predictive modeling with Mills-ratio-based correction to enable the estimation of censored values in a compositionally coherent manner. The framework is demonstrated using different ML techniques, including random forest (RF) and extreme gradient boosting (XGB), on geochemical datasets from the Geological Survey of Norway, comprising both artificially censored copper data with known ground truth and naturally censored silver, mercury, and tantalum data. Results show that ReX-CoDA performs competitively against established imputation methods and provides a robust, scalable, and compositionally consistent alternative for handling censored geochemical data.