Targeted maximum likelihood estimation for psychological research: From causal identification to statistical inference.
Causal questions have long been central to psychological research, particularly in randomized experiments, while formal causal-inference methods are increasingly being applied to observational and quasi-experimental data. Common outcome-regression and propensity-score approaches can be sensitive to nuisance-model misspecification, whereas flexible machine learning alone does not ensure valid target-parameter estimation or inference. Targeted maximum likelihood estimation (TMLE) addresses this estimation problem by combining outcome and treatment information through an efficient-influence-function-based update, yielding doubly robust estimation and influence-function-based inference under appropriate conditions. TMLE targets a parameter of the observed-data distribution; causal interpretation additionally requires a well-defined intervention, causal model, and identification assumptions. Although the Causal Roadmap and targeted-learning literature establish the general framework, guidance for modern applied implementation remains fragmented. This research integrates causal-to-statistical parameter mapping, accessible semiparametric explanation, reproducible Python code, positivity diagnostics, fully nested cross-fitting, and TMLE-specific reporting guidance for psychological researchers. The formal exposition and simulations use the point-treatment average treatment effect for a binary treatment or exposure as an illustrative target. The two empirical examples are instead framed as statistical implementation exercises targeting covariate-standardized observed-data risk differences, with the assumptions under which those parameters would equal causal effects stated separately. The examples contrast transparent parametric estimation with cross-fitted Super Learner TMLE, while the simulations examine nuisance-model misspecification, limited overlap, propensity-score truncation, and flexible TMLE with and without cross-fitting. The contribution is translational rather than theoretical: TMLE is presented as the estimation stage of a broader causal-inference workflow, not as a procedure that independently establishes causality.