A Multi-Dimensional Benchmark of HTE Estimators for Decision Support
Abstract
Estimating heterogeneous treatment effects (HTEs) underlies personalized decision-making in domains ranging from precision medicine to targeted marketing, yet estimator behavior under large samples, high-dimensional nuisance covariates, and non-random treatment assignment remains incompletely characterized. We benchmark seven representative HTE estimators: S-, T-, and X-Learner; LinearDML and LinearDRLearner; and CausalForestDML and ForestDRLearner. The large-scale experiments use standardized empirical covariates from the Criteo Uplift dataset with simulated treatment and outcomes, scale to 106 observations, and include linear and nonlinear CATE functions. Under these evaluated semi-synthetic settings, linear orthogonal learners are most accurate when the CATE is linear but plateau under nonlinear heterogeneity, whereas forest-based variants are more adaptive at greater computational cost. A scenario-wise descriptive ranking summarizes the observed accuracy–robustness–cost trade-offs without asserting a universally best estimator.