Aug 2026· 2026 12th International Conference on Big Data and Information Analytics (BigDIA)· pp. 621-628· 0 citations· 19 references
Abstract
Estimating heterogeneous treatment effects (HTEs) underlies personalized decision-making in domains ranging from precision medicine to targeted marketing, yet estimator behavior under large samples, high-dimensional nuisance covariates, and non-random treatment assignment remains incompletely characterized. We benchmark seven representative HTE estimators: S-, T-, and X-Learner; LinearDML and LinearDRLearner; and CausalForestDML and ForestDRLearner. The large-scale experiments use standardized empirical covariates from the Criteo Uplift dataset with simulated treatment and outcomes, scale to 106 observations, and include linear and nonlinear CATE functions. Under these evaluated semi-synthetic settings, linear orthogonal learners are most accurate when the CATE is linear but plateau under nonlinear heterogeneity, whereas forest-based variants are more adaptive at greater computational cost. A scenario-wise descriptive ranking summarizes the observed accuracy–robustness–cost trade-offs without asserting a universally best estimator.
Biological research and clinical evidence suggest that treatment response may vary substantially along characteristics, such as comorbidities, genetic variants, environmental, or socio-economic factors. Future precision medicine requires accurate assessment of heterogeneous treatment effects (HTE) to guide optimal clin...
This work analyzes current benchmarking practices and introduces a novel decomposition framework that disentangles the contribution of distinct data-generating components, such as confounding, dose distribution non-uniformity, and response surface complexity, to estimator performance.
Christopher Bockel-Rickermann, Daan Caljon, Toon Vanderschueren et al.· Proceedings of the 32nd ACM...· 2 citations
Doubly robust survival estimators require nuisance models that the available information can support. We examine design and estimator dependence in TCGA breast cancer data. Treatment indicators disagree for 337 of 1076 patients; 67.4% have unobserved five-year binary outcomes; and receptor imputation expands the triple...
W. Fahmy, E. Krikun, A. Al Khateeb· medRxiv· 0 citations
A new inference method for conducting multiple-treatment comparisons involving endpoints within the generalized linear model (GLM) framework under covariate-adaptive randomization (CAR) that can effectively control Type I error while potentially improving power.
In various fields, such as medicine and marketing, accurately predicting individual treatment effects holds significant promise. However, achieving reliable predictions alone is often insufficient for making informed decisions; it is equally important to understand why the treatment effect is higher for some individual...
Theoretically, it is proved that when the auxiliary outcomes satisfy a set of surrogacy conditions and the representation retains relevant covariate information, the original CATE is identified when the high-dimensional covariates are replaced by the learned representation.
Maitreyi Swaroop, Shikha Bhat, Samantha Rodriguez et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.