Skip to content
Preprint

Separating Covariate Shift from Mechanism Change with Two Discriminators: CJSD, a Conditional Discrepancy with an Exact Covariate-Concept Decomposition

Aug 2026 · 1 citation · 8 references
Computer Science

TL;DR

The Conditional Jensen-Shannon Discrepancy (CJSD) is proposed, which proves a covariate-null property, a drift-mass law, a one-sided misspecification-control inequality, a one-sided misspecification-control inequality, and a fixed-measure metrization via an identifiability lemma.

Abstract

After the inputs X are known, how much additional information does the label Y carry about which dataset a sample came from? That single quantity -- estimable as the difference of two discriminators'held-out cross-entropies, D_CJS = CE(Z|X) - CE(Z|X,Y) -- is exactly the part of a dataset difference that covariate shift cannot explain. We propose the Conditional Jensen-Shannon Discrepancy (CJSD): with a task indicator Z, the chain rule I(Z;X,Y) = I(Z;X) + I(Z;Y|X) splits total task discrepancy exactly into a covariate axis and a functional axis, both estimable from two ordinary classifiers, with no task-specific predictors, generative models, or bootstrap surrogates. We prove a covariate-null property (the functional axis is exactly zero under pure covariate shift, however severe), a drift-mass law (D_CJS/ln2 equals the mass of the disagreement region for deterministic labels), a one-sided misspecification-control inequality (each direction of estimation error is bounded, unconditionally, by the excess risk of a single discriminator), and a fixed-measure metrization via an identifiability lemma. Empirically, on a ten-measure battery over 202 dataset pairs (synthetic, Electricity, Covertype), only the two conditional-information estimators -- CJSD and a kNN plug-in for the same estimand -- separate concept from covariate shift with AUC 1.0; the case for CJSD is the estimator: under controlled dimensionality scaling the kNN plug-in fails from d=64 while the discriminator route holds to d=256 with a swappable classifier, and it alone yields paired confidence intervals and sequential extensions from the same learned object. The same estimator audits the conditional fidelity of synthetic-data generators that marginal and joint QA metrics pass, detects annotation-guideline changes invisible to input-space monitors, and supports null-calibrated fairness audits.

View source

Similar papers

Jul 2026

Adaptive deep nonparametric regression from dependent data under covariate shift

This paper considers deep neural network estimators for nonparametric quantile and Huber regression under covariate shift and from dependent observations and proposes a sparse-penalized deep neural network (SPDNN) estimator that takes into account the discrepancy between the source and target distributions of the data.

W. Kengne, Ehud Mossa Ockegna · 0 citations
Preprint Aug 2026

Coarsening Latent-Class Probabilities: Directional Distortion and Coverage Loss

Outcomes are regressed on a calibrated probability vector for unobserved class membership. Under a structural conditional mean excluding the score and conditional calibration, the observed-data model reduces to a partially linear regression. The probability vector is a Berkson-type surrogate for membership, so the effe...

M. T. Kurbucz · 0 citations
Preprint Sep 2026

Optimizer as Detector: Stochastic Gradient Descent for Latent Mixture Models

Pooling latent subpopulations can obscure relationships and yield misleading regression conclusions, including Simpson's paradox (SP). We propose a detector based on the steady-state dynamics of constant-step stochastic gradient descent (SGD). Unlike likelihood-based mixture tests and confounder-search methods, it requ...

Ye Shi, Xiao Jin, Chung-Piaw Teo · 0 citations
Preprint Jul 2026

Bridging extrinsic and intrinsic variable importance

Simulations show that VIMP and MPLOCO agree most closely when the fitted learner is well aligned with the data-generating mechanism, and clarify when intrinsic and extrinsic importance can be interpreted similarly and when they provide complementary information.

Yucheng Zhao, Brian D. Williamson · 0 citations
Preprint Aug 2026

Inferential Evaluation of Surrogate-Derived Models under Covariate Shift

Cross-fitted estimators that transport information from the two labeled sources through source-specific density ratios are proposed that establish asymptotically linear inference for TPR and FPR, consistency and pointwise inference for the ROC curve, and asymptotically normal inference for AUC.

Long-Tian Shi, Molei Liu, Dou-Dou Zhou · 0 citations
#artificial intelligence Preprint Sep 2026

General Quantification of Covariate and Concept Shifts

This paper proposes a key notion: $\gamma^{*}\!$-concept shifts, and derive a general error bound unifying covariate and $\gamma^{*}\!$-concept shifts, which applies to broad loss functions, label spaces, and stochastic labeling and develops estimators for these shifts with concentration guarantees.

Hong-Bo Chen, L. Xia · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.