Skip to content
Preprint

Identification and Inference with Machine-Learned Instruments

Jul 2026 · 0 citations
Economics

TL;DR

A heterogeneity-robust orthogonal score is constructed that restores $\sqrt{N}$ inference on the fixed, learner-invariant target at no efficiency cost, and a Hausman-type diagnostic and identification-robust confidence sets are provided.

Abstract

Instrumental-variables estimation increasingly pools many or high-dimensional instruments into a single machine-learned first stage, with rich controls partialled out. The resulting estimand, the partialled-out IV coefficient built from any signal of the instruments, is a signal-weighted average of the heterogeneous effects, which gives an opaque first stage a precise structural meaning. The average is convex whenever a covariance-monotonicity condition holds, and we provide a microfoundation for that condition based on vector monotonicity. With a learned signal, however, the usual debiased moment is not Neyman-orthogonal, and its first-order bias is a drift toward the learner's own signal-weighted average, so naive inference remains valid only for that learner-dependent target. We construct a heterogeneity-robust orthogonal score that restores $\sqrt{N}$ inference on the fixed, learner-invariant target at no efficiency cost, and provide a Hausman-type diagnostic and identification-robust confidence sets.

View source

Similar papers

Preprint Jul 2026

Best-Arm Identification with Generative Proxy

Best-arm identification is a canonical model for data-driven decision-making, but in many applications each reward observation is costly. Motivated by the growing availability of cheap predictions from machine learning and large language models, we study fixed-confidence best-arm identification in which each costly reward pull is paired with a cheap but correlated proxy score. The marginal mean of the proxy can be estimated offline and is treated as known, whereas its correlation $\rho$ with the reward, which governs how much the proxy helps, is unknown and must be learned online in pair with real rewards. We show that a control-variate adjustment turns this model into a heteroscedastic identification problem whose oracle sample complexity improves by residual variance $1-\rho^2$. The central difficulty is that the correlation must be learned from the same costly samples that identification consumes online, and that a plug-in estimate of the residual variance is anti-conservative and can compromise correctness. We propose PROBE (PRoxy OLS for Best-arm Exploration), a phase-elimination algorithm that directly maintains an upper certificate on the residual variance with an ordinary least squares fit, whose exact chi-square law keeps the certificate valid regardless of the unknown correlation. We prove that PROBE is $\delta$-PAC and attains the known-correlation oracle sample complexity up to a constant multiplicative factor and a constant additive calibration cost. The guarantee extends to the $(\epsilon,\delta)$-PAC setting under minimal changes to the algorithm. Numerical experiments on synthetic instances and on an auto-loan pricing replay with large language model and tabular proxies confirm that the sample savings of PROBE scale with the strength of the reward-proxy correlation, exactly as the theory predicts.

Tianyi Ma, Hanzhang Qin, Ruihao Zhu et al. · 1 citation
Preprint Aug 2026

Double/Debiased Machine Learning for Functional-Form-Robust Spatial Autoregression

Spatial autoregressive inference is typically conditional on the spatial weights matrix, W, even though the underlying interaction structure is often unknown and empirical conclusions can be sensitive to its specification. This paper develops double/debiased machine learning inference for low-dimensional SAR parameters when the spatial interaction operator is learned flexibly from potentially endogenous characteristics. Within a maintained admissible support, interaction strength is generated by an unknown function of geographic and socioeconomic characteristics, making inference robust to functional form specification of the weights within that support. Endogeneity in the characteristics generating W is addressed through a nonlinear control function based on locally relevant first-stage residual information. Because the learned operator enters both the spatial lag and spatially transformed instruments, treating the estimated W as known generally leaves a first-order generated-W effect. I construct an operator-orthogonal SAR-IV/GMM score that removes this leading sensitivity and combine it with buffered spatial cross-fitting that separates evaluation-score footprints from nuisance-training observations. Under near epoch dependence on a spatially mixing innovation field and target-relevant nuisance rate and regularity conditions, the estimator is asymptotically linear and root-n normal. Monte Carlo simulations show improved finite-sample inference relative to nonorthogonal alternatives when the interaction function is misspecified, weight generating characteristics are endogenous, and observations are spatially dependent. In a U.S. application, diabetes estimates vary with the choice of W, showing the sensitivity of SAR inference to the interaction structure. Even for the same learned W, results differ across inferential methods, highlighting the importance of inference when W is learned.

Jieun Lee · 0 citations
Preprint Jul 2026

A Machine-Learning-Compatible Omnibus Test for Treatment Effect Heterogeneity

This study proposes a formal, computationally efficient nonparametric omnibus test for treatment-effect heterogeneity that is compatible with a broad class of estimators, including modern machine-learning methods and is illustrated using two empirical applications on retirement savings and trade liberalization.

Elia Lapenta, Anthony Strittmatter, Pedro Vergara Merino · 0 citations
Open access Aug 2026

Correct (and Incorrect) Inference with a Single Instrumental Variable: Practical Takeaways from the Weak Instruments Literature

Most empirical economists have encountered the warning that instrumental variables can be “weak,” but the underlying issues—what makes an instrument weak, why weakness distorts inference, and what to do about it—are less widely understood. This article offers an accessible introduction to the weak instruments problem for the common just-identified case of a single endogenous regressor and a single instrument. We explain why the usual two-stage least squares t-ratio and its “±1.96 times the standard error” confidence interval can yield incorrect inferences, much as homoskedasticity-only standard errors do when errors are not homoskedastic. We then describe practical, robust-to-weak-instrument solutions—including the Anderson-Rubin and tF methods—that deliver valid confidence intervals whatever the instrument's true strength, and we offer some do's and don'ts, notably why the popular “F greater than 10” rule has no theoretical justification in this setting.

David S. Lee, Jack R. Porter · 0 citations
Preprint Aug 2026

Exact Inference in Fixed-Effect Regressions with Concentrated Identifying Variation

In fixed-effect regressions with many groups, fixed effects can absorb most identifying variation, leaving a handful of observations to carry what remains. When variation is this concentrated, conventional $t$-tests can reject a true null more than half the time, and any fixed critical value is either invalid or so conservative it has essentially no power. This paper builds an exact test from the design alone. A \textit{nuisance-annihilating contrast} is a linear combination of the treatment and fixed-effect dummies that eliminates the fixed effects without touching the outcome; sign-flipping these contrasts is then an exact symmetry of the null distribution at every sample size, under arbitrary heteroskedasticity. In two-way designs --- worker-firm, firm-time --- these contrasts are exactly the cycles of the bipartite mobility graph, so the movement that identifies the treatment effect is what makes exact inference possible. Exactness costs power: relative to an oracle test, a chosen set of cycles has an observable \textit{capture ratio} $\kap\in[0,1]$ and standard-error premium $\kap^{-1/2}$, and a packing algorithm resolves the capture-granularity trade-off. In the Grunfeld investment regression (single-observation score concentration $73.9\%$), 32 cycle contrasts capture $\kap=0.627$ of the identifying variation, giving an exact $95\%$ confidence interval of $[0.150,\,0.450]$.

S. Halkiewicz · 0 citations
Preprint Jul 2026

Identification and Robust Inference for Multiple Treatment Effects with Possibly Invalid Instruments

The instrumental variable (IV) method is widely used to infer causal effects in observational studies with unmeasured confounding, but invalid instruments can compromise both population identification and finite-sample inference. This paper studies linear IV models with multiple endogenous treatments and possibly invalid instruments. Identification is more delicate than in the single-treatment setting because a single instrument no longer identifies a scalar candidate effect; instead, each relevant instrument defines a hyperplane in the multidimensional effect space. For identification of multiple treatment effects, we introduce generalized plurality and majority rules which require a sufficiently large number of IVs to be valid. For inference, data-dependent instrument selection may fail to separate certain invalid IVs from valid ones, leading to undercoverage of confidence intervals when these invalid instruments are mistakenly selected as valid. We propose a sampling confidence interval for each treatment effect, which is robust to IV selection errors. We establish asymptotic coverage and parametric-rate length of our sampling confidence interval under regularity conditions and illustrate this method in a Mendelian randomization application.

Ziwei Mei, Qingliang Fan, Zijian Guo · 0 citations