Skip to content
Preprint

A Machine-Learning-Compatible Omnibus Test for Treatment Effect Heterogeneity

Jul 2026 · 0 citations · 40 references
Economics

TL;DR

This study proposes a formal, computationally efficient nonparametric omnibus test for treatment-effect heterogeneity that is compatible with a broad class of estimators, including modern machine-learning methods and is illustrated using two empirical applications on retirement savings and trade liberalization.

Abstract

This study proposes a formal, computationally efficient nonparametric omnibus test for treatment-effect heterogeneity that is compatible with a broad class of estimators, including modern machine-learning methods. The test is designed for settings in which identification can rely on high-dimensional controls while heterogeneity is assessed with respect to a low-dimensional subset of covariates. We derive the test statistic's asymptotic null distribution and develop a bootstrap procedure that is efficient because it avoids re-estimating nuisance parameters in each iteration. The testing approach applies to multiple empirical designs, including randomized experiments, selection-on-observables, difference-in-differences, and instrumental-variables settings. Monte Carlo simulations show that the test attains near-nominal size under the null and exhibits good power against heterogeneous alternatives. We further illustrate the procedure using two empirical applications on retirement savings and trade liberalization.

View source

Similar papers

Preprint Jul 2026

Orthogonal Integrated Conditional Moment Tests for Treatment Effect Heterogeneity

We propose a nonparametric integrated conditional moment (ICM) test for treatment effect heterogeneity across subpopulations defined by a given covariate subvector. Under unconfoundedness, the null is recast as a conditional moment restriction based on a Neyman-orthogonal score, which reduces the first-order sensitivity of the empirical process to nuisance parameter estimation. The test statistics are constructed as continuous functionals of a marked empirical process. We establish a uniform feasible-to-oracle approximation and derive the asymptotic properties of these test statistics under the null and fixed alternatives. We further show that the test has nontrivial power against local alternatives converging to the null at the $n^{-1/2}$ rate, and develop an easy-to-implement multiplier bootstrap for feasible inference. We also develop extensions to tests of parametric CATE specifications and to settings with endogenous treatment and a binary instrument. Finally, we apply the proposed testing approach to study whether the effect of maternal smoking during pregnancy on infant birth weight varies with maternal age.

Hao Lu, Xiaojun Song · 0 citations
Preprint Jul 2026

Testing the equality of estimable parameters

This paper proposes a general and unified framework for testing the equality of a broad class of parameters, defined via $U$-statistics, across multiple independent populations. This approach encompasses various common statistical problems, such as comparing variances, correlation coefficients, or Gini indices, among many others. We consider two test statistics, a Wald-type statistic and an ANOVA-type statistic. The asymptotic distribution of the first one is derived under a fixed-dimension regime, whereas the second one is studied under both fixed and increasing-dimension regimes, where the parameter dimension diverges with the sample size. Based on these limiting distributions, we construct test procedures enabling asymptotically exact inference without parametric assumptions. Additionally, an alternative null distribution estimator based on a weighted bootstrap approximation is studied, which is applicable to the ANOVA-type statistic under a fixed-dimension regime. The finite-sample performance and computational efficiency of the proposed procedures are evaluated through an extensive simulation study. Finally, an application to a real dataset illustrates the usefulness of the proposed methodology.

M. Romero-Madronal, M. R. Sillero-Denamiel, M. D. Jiménez–Gamero · 0 citations
Preprint Jul 2026

A Variance-Based Test for Heterogeneous Treatment Effects

This paper proposes a robust nonparametric hypothesis test for the existence of heterogeneous treatment effects. We focus on the variance of the Conditional Average Treatment Effect (CATE) as a natural omnibus parameter, where a non-zero variance implies the presence of relevant heterogeneity. Standard inference for this parameter faces a fundamental theoretical challenge. On one hand, evaluating variance components on the same sample leads to null degeneracy, where the asymptotic variance collapses to zero under the null hypothesis of homogeneity, invalidating standard Gaussian inference. On the other hand, decoupling the empirical processes via standard sample-splitting breaks the Neyman orthogonality of the doubly robust scores due to their nonlinear squared loss, which prevents the cancellation of first-order regularization biases. To resolve this challenge, we propose a novel Intra-Fold Sample-Splitting algorithm. By evaluating variance components on mutually disjoint subsamples while coupling them to identical out-of-fold nuisance estimators, our procedure achieves algebraic cancellation of the nuisance biases. We prove this restores consistency and asymptotic normality, and ensures Type I error control. Monte Carlo simulations demonstrate that the proposed test achieves superior size control relative to existing tests while maintaining high power. In an empirical application to the NSW job training program, the test detects significant heterogeneity that traditional nonparametric tests fail to uncover.

Fang Yu · 0 citations
Preprint Aug 2026

Double Machine Learning with High-dimensional Interactive Fixed Effects

Factor structures are central to empirical work in economics and finance, and are usually used to model time-varying unobserved heterogeneity through interactive fixed effects (IFE). Existing IFE estimators rest on low-dimensional and linear specifications in the covariates, assumptions which are increasingly restrictive in applications drawing on rich datasets with controls of unknown functional form. This paper develops a Double Machine Learning estimator for the high-dimensional partially linear panel model with interactive fixed effects (panel DML-IFE). The method combines projection-based defactorisation of the data, in the spirit of Common Correlated Effects (CCE), with a Neyman-orthogonal score function and cross-fitting procedure, and accommodates low-rank factor structures in outcomes and treatments alongside high-dimensional, potentially nonlinear covariate effects estimated by machine learning algorithms. Monte Carlo simulations show that panel DML-IFE outperforms conventional IFE estimator outside the correctly-specified linear case, with bias reduction driven primarily by the time and covariate dimensions. An empirical application to U.S. stock returns shows that several effects documented under linear specifications lose statistical significance once high-dimensional nonlinear confounding and the presence of IFE are jointly accounted for.

Bin Chen, Annalivia Polselli, P. Clarke · 0 citations
Preprint Aug 2026

Fast high-dimensional mean testing via logistic regression

We propose computationally efficient tests for equality of mean vectors of two or more high-dimensional populations. Central to our approach is an equivalence between equality of means and a zero population logistic regression parameter. We establish this equivalence for independently distributed observations without imposing common distributional assumptions across populations. Our procedure uses logistic Lasso to screen informative variables and an unpenalized logistic refit for inference in the reduced dimension, yielding asymptotically correct size and consistency. For a specified two-sample Gaussian submodel and sparse discriminative class, the test also attains the minimax separation rate. The framework extends to multiple populations through multi-class logistic regression. Simulations demonstrate accurate size control, strong power, and favorable computational scaling compared with existing tests under unbalanced designs and variance heterogeneity. Applications to gene-expression data with more than twenty-two thousand variables illustrate the practical scalability of the proposed procedures.

Sayan Das, Debraj Das, S. Dutta · 0 citations
Open access Aug 2026

An adaptive test for two-sample accelerated life models

Abstract We propose an adaptive testing method that is robust for two-sample scale models with censored observations. Motivated by [H. Uno, L. Tian, B. Claggett and L. J. Wei, A versatile test for equality of two survival functions based on weighted differences of Kaplan–Meier curves, Stat. Med. 34 2015, 28, 3680–3695], we propose simulation-based procedures to check model validity that exhibit robust performance across a broad range of alternative hypotheses. To evaluate the behavior of the proposed test, we conduct comprehensive simulations in some widely used survival functions. Simulation results indicate that the test exhibits strong performance in detecting scale difference between two samples, demonstrating adequate power. The proposed procedures are illustrated using a real-world dataset.

Seung-Hwan Lee · 0 citations