This work studies a different regime in which the probability of label missingness depends on posterior classification uncertainty, so that the observed missing-label indicators can themselves carry information about the Bayes decision boundary.
Abstract
Semi-supervised classifiers are commonly trained from samples in which all features are observed but some class labels are missing. When label missingness is independent of the observed data, unavailable class memberships reduce Fisher information relative to a completely classified sample. We study a different regime in which the probability of label missingness depends on posterior classification uncertainty, so that the observed missing-label indicators can themselves carry information about the Bayes decision boundary. Building on the conditionally weighted information decomposition of Ahfock and McLachlan, we develop this phenomenon for a two-component exponential mixture. Although the exponential model is non-Gaussian, asymmetric, and supported on the positive half-line, its log-posterior odds remain linear in the feature. We derive Bayes'rule and its exact error rate, formulate entropy-logistic and squared-discriminant missingness mechanisms, and obtain the full partially classified likelihood. We then derive a decomposition of the Fisher information into the complete-data information, the conditionally weighted loss due to missing labels, and the information contributed by the missing labels. Numerical quadrature identifies regions in which the full likelihood classifier has asymptotic relative efficiency above or below one. Monte Carlo experiments with finite training samples broadly support the population calculations, with the largest departures from the asymptotic predictions occurring near the transition at which the relative efficiency crosses one.
Numerical studies and a semi-synthetic analysis based on hard-drive failure data illustrate potential reductions in expected error rate and improvements in decision-boundary estimation from modelling feature-dependent label missingness.
Jinran Wu, You‐Gan Wang, Geoffrey J. McLachlan· 0 citations
Gaussian-mixture calculations and a medical diagnosis example illustrate how uncertainty-dependent labeling mechanisms can improve estimation and classification under a fixed labeling budget.
You‐Gan Wang, Jin-Ran Wu, Geoffrey J. McLachlan· 0 citations
A classification-weighted generalized-eigenvalue criterion is developed under which informative partial classification may have smaller asymptotic classification risk without globally dominating complete classification in Fisher information.
Fariborz Setoudehtazang, Geoffrey J. McLachlan· 0 citations
Classification without Labels (CWoLa) shows that, in the binary case, a classifier trained to distinguish two impure mixtures with different class proportions can recover an optimal class discriminator without knowing the mixture proportions, and proposes prior-free procedures that train a standard classifier to distin...
Raphaël Bonnet-Guerrini, Johann Ioannou-Nikolaides, Troels C. Petersen et al.· arXiv.org· 1 citation
We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk. We consider parametric classification problems in which features are observed for all sampling units while class labels can be acquired selectively. By combining the Fisher information supplied by an acquired...
F. Setoudehtanzangi, Geoffrey J. McLachlan· 0 citations
Incomplete compositional data analysis faces a fundamental limitation: likelihood-based methods for compositional models generally require fully observed compositions, making it difficult to accommodate missing or censored proportions directly on the simplex. Consequently, analysts often discard partially observed comp...
J. Pillay, A. Bekker, C. Tortora et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.