Finite-sample conditional tests of a common success probability within a parent selected by deterministic Gini CART, which separate information loss from computational limitations and exhibit a selected fiber disconnected under single-label swaps.
Abstract
Binary classification trees select subgroups using the same outcomes later used to assess their differences. We develop finite-sample conditional tests of a common success probability within a parent selected by deterministic Gini CART. The construction retains all eligible cutpoints and conditions on the selected split, its ancestor path, the parent success total, and outside outcomes. The resulting uniform label fiber gives an exact count distribution, while a reversible parallel Monte Carlo construction yields super-uniform inclusive and exactly uniform tie-randomized p-values for any prespecified finite run budget. Within a two-child constant-risk model, the selected count law is an exponential family with information equal to its conditional count variance. We separate information loss from computational limitations and exhibit a selected fiber disconnected under single-label swaps. Simulations with 200 or 400 observations, ten independent or correlated predictors, and trees of depth three show conservative inclusive tests and nontrivial power for large risk differences. At 400 observations and a generating risk difference of 0.4, randomized rejection conditional on reaching the prespecified third-level target is 51--71\%, while selection followed by rejection occurs in 11--21\% of datasets. Smaller signals remain difficult to detect, and a representative fivefold increase in computation gives little power improvement. The guarantee concerns parent homogeneity, or equality of two constant child risks, and does not cover equality of heterogeneous regional averages.
This paper proposes Wasserstein Causal Forests (WCF) for settings in which each unit's outcome is itself a probability distribution. This study also defines finite-grid transformed average and conditional average treatment effects, including a reference-distance contrast that asks whether treatment moves unit-level dis...
Comparing two populations at the same physical covariate value requires more than conditional means or isolated target-point decisions: researchers may need evidence about an entire conditional-distribution ordering over a continuum, even when covariate margins differ. This paper makes that common-value comparison esti...
We give an exact randomization-based confidence set for the average treatment effect (ATE) in matched-pair studies with a binary outcome, requiring neither monotonicity nor any distributional assumption beyond the within-pair coin flip. At its core is an analytic solution to the worst-case allocation of attributable ef...
In various fields, such as medicine and marketing, accurately predicting individual treatment effects holds significant promise. However, achieving reliable predictions alone is often insufficient for making informed decisions; it is equally important to understand why the treatment effect is higher for some individual...
Statistical validation of synthetic multivariate data requires assessing whether a generator preserves the joint dependence structure of the target population without merely reproducing observed records. We develop a model-agnostic framework based on full conditional distributions. For each coordinate, we normalize the...
Reference-based multiple imputation is used in longitudinal clinical trials to assess sensitivity to assumptions about outcomes unobserved after intercurrent events. Most existing methods target continuous outcomes and use multivariate normal working models. Scheduled count outcomes require a model that preserves integ...
D. Burger, Emmanuel Lesaffre, R. Martina· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.