Model-assisted estimation with a training subsample: a two-phase sampling approach with design-based variance estimation
Abstract
When a flexible prediction model is fitted on a training subsample drawn from a probability sample, the model-assisted estimator actually reported arises from one realized partition, yet existing theory quantifies uncertainty only for partition-averaged, cross-fitted, or symmetrized versions of it. We represent the training subsample as a second phase of sampling and derive, exactly and for any algorithm, a two-term variance decomposition and the variance family linking the single-partition estimator to its Rao-Blackwellized average, whose design bias it shares. For tree-type predictors the second-phase variance is computable in closed form when the cell structure is fixed or s-measurable, and its share of total variance grows with tree complexity, contributing to documented variance underestimation through a mechanism distinct from residual shrinkage. We propose an analytic and a replication variance estimator, neither altering the point estimate, and evaluate them by simulation: in the populations studied, budgeting the second phase recovers most of the coverage lost by ignoring it, at a small fraction of the cost of partition averaging.