Jun 2024· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· 2 citations· 79 references
Computer Science
TL;DR
This work analyzes current benchmarking practices and introduces a novel decomposition framework that disentangles the contribution of distinct data-generating components, such as confounding, dose distribution non-uniformity, and response surface complexity, to estimator performance.
Abstract
Standard supervised learning algorithms prioritize predictive performance over causal inference, making them ill-suited for conditional average dose response (CADR) estimation. Although specialized CADR estimators have been proposed to address this, the field lacks clarity on which dataset characteristics truly challenge model performance, largely because current benchmarking practices fail to isolate different sources of estimation error. In this work, we analyze these current benchmarking practices and introduce a novel decomposition framework that disentangles the contribution of distinct data-generating components, such as confounding, dose distribution non-uniformity, and response surface complexity, to estimator performance. Applying this scheme to several established benchmarks, we uncover that widely used datasets do not primarily test for confounding robustness, as often assumed, but are instead dominated by challenges arising from non-uniform dose distributions. We further propose a new benchmark dataset with high CADR heterogeneity, where confounding does have a substantial impact. Our results call for a rethinking of current evaluation practices and advocate for more diagnostic, data-centric benchmarks to advance the development of robust CADR estimation methods.
Self-supervised Causal Effects Estimation is proposed, a novel framework that integrates causal priors with self-supervised learning to construct balanced and predictive representations for causal effects estimation that consistently outperforms state-of-the-art methods.
Xin-Shu Li, Shiyi Yang, Venus Haghighi et al.· ACM Transactions on Intellig...· 0 citations
Estimating heterogeneous treatment effects (HTEs) underlies personalized decision-making in domains ranging from precision medicine to targeted marketing, yet estimator behavior under large samples, high-dimensional nuisance covariates, and non-random treatment assignment remains incompletely characterized. We benchmar...
Ya-Xin Zhang, Quan-Dong Wang, Xiao-Min Zhu et al.· 2026 12th International Conf...· 0 citations
Estimating conditional average treatment effects (CATE) enables efficient targeting of interventions, but many applications have limited experimental samples, making it difficult to estimate heterogeneous effects from high-dimensional covariates. In such settings, policymakers and medical practitioners often succumb to...
Maitreyi Swaroop, Shikha Bhat, Samantha Rodriguez et al.· 0 citations
High-dimensional data create challenges for causal effect estimation because identifying the covariates needed for correct model specification becomes increasingly difficult. Double/debiased machine learning (DML) facilitates the use of machine learning (ML) for causal inference by mitigating regularization and overfit...
Studying heterogeneous treatment effects has become essential in experimental and observational studies. A critical assumption for obtaining reliable treatment effect estimates is overlap, which requires that treated and control units have sufficiently similar covariate distributions. Poor overlap may limit the effecti...
Post-click conversion rate (CVR) is a key metric in various scenarios including e-commerce and advertising, reflecting the efficiency and user experience in the second stage of the conversion process. Estimating the causal effect on CVR is therefore of great practical importance. However, directly applying existing cau...
Jia-Yi Dan, Bo Li, Lu Deng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.