Skip to content
Conference

A Multi-Dimensional Benchmark of HTE Estimators for Decision Support

Aug 2026 · 2026 12th International Conference on Big Data and Information Analytics (BigDIA) · pp. 621-628 · 0 citations · 19 references

Abstract

Estimating heterogeneous treatment effects (HTEs) underlies personalized decision-making in domains ranging from precision medicine to targeted marketing, yet estimator behavior under large samples, high-dimensional nuisance covariates, and non-random treatment assignment remains incompletely characterized. We benchmark seven representative HTE estimators: S-, T-, and X-Learner; LinearDML and LinearDRLearner; and CausalForestDML and ForestDRLearner. The large-scale experiments use standardized empirical covariates from the Criteo Uplift dataset with simulated treatment and outcomes, scale to 106 observations, and include linear and nonlinear CATE functions. Under these evaluated semi-synthetic settings, linear orthogonal learners are most accurate when the CATE is linear but plateau under nonlinear heterogeneity, whereas forest-based variants are more adaptive at greater computational cost. A scenario-wise descriptive ranking summarizes the observed accuracy–robustness–cost trade-offs without asserting a universally best estimator.

View source

Similar papers

Preprint Aug 2026

Martingale R-learner: Estimating Time-varying Heterogeneous Treatment Effects for Time-to-event Outcomes

Biological research and clinical evidence suggest that treatment response may vary substantially along characteristics, such as comorbidities, genetic variants, environmental, or socio-economic factors. Future precision medicine requires accurate assessment of heterogeneous treatment effects (HTE) to guide optimal clin...

Jue Hou, Yuchen Qi, Rong-Hui Xu · 0 citations
Book Open access Jun 2024

A Data-Centric Decomposition of Estimator Performance in Continuous Treatment Effect Estimation

This work analyzes current benchmarking practices and introduces a novel decomposition framework that disentangles the contribution of distinct data-generating components, such as confounding, dose distribution non-uniformity, and response surface complexity, to estimator performance.

Christopher Bockel-Rickermann, Daan Caljon, Toon Vanderschueren et al. · 2 citations
Open access Oct 2026

Design and outcome-model estimability in doubly robust survival analysis: a breast cancer reanalysis

Doubly robust survival estimators require nuisance models that the available information can support. We examine design and estimator dependence in TCGA breast cancer data. Treatment indicators disagree for 337 of 1076 patients; 67.4% have unobserved five-year binary outcomes; and receptor imputation expands the triple...

W. Fahmy, E. Krikun, A. Al Khateeb · 0 citations
Open access Aug 2026

Valid Test for Multi-arm Trials with Generalized Linear Models Under Covariate-adaptive Randomization

A new inference method for conducting multiple-treatment comparisons involving endpoints within the generalized linear model (GLM) framework under covariate-adaptive randomization (CAR) that can effectively control Type I error while potentially improving power.

Guannan Zhai, Feifang Hu · 0 citations
#machine learning Preprint Sep 2026

Splitting the Difference: Interpretable Causal Forests for Treatment Effect Heterogeneity and Bias

In various fields, such as medicine and marketing, accurately predicting individual treatment effects holds significant promise. However, achieving reliable predictions alone is often insufficient for making informed decisions; it is equally important to understand why the treatment effect is higher for some individual...

N. Ihlo, Merle Behr · 0 citations
#machine learning Preprint Sep 2026

Representation Learning for Sample-Efficient CATE Estimation by Leveraging Multiple Outcomes

Theoretically, it is proved that when the auxiliary outcomes satisfy a set of surrogacy conditions and the representation retains relevant covariate information, the original CATE is identified when the high-dimensional covariates are replaced by the learned representation.

Maitreyi Swaroop, Shikha Bhat, Samantha Rodriguez et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.