Skip to content
Open access

Causal Benefit-Aware Recommendation for Personalized Learning-Path Features: A Targeting-Policy Framework with Provable Guarantees and Randomized Evaluation

Aug 2026 · Electronics · Vol 15, pp. 3483 · 0 citations · 32 references

TL;DR

This work formalizes feature recommendation as a causal targeting-policy problem: rank students by the estimated conditional average treatment effect (CATE) of a feature and recommend to the top of the ranking, proving causal top-CATE targeting maximizes policy value at any budget and weakly dominates predictive (outcome-based) targeting.

Abstract

Educational platforms increasingly personalize which AI learning-path features (adaptive homework, learner choice) each student receives. The natural correlational baseline ranks students by predicted performance—deliver the feature to those expected to do well—a heuristic that need not identify who actually benefits. We formalize feature recommendation as a causal targeting-policy problem: rank students by the estimated conditional average treatment effect (CATE) of a feature and recommend to the top of the ranking. We prove three results: (i) causal top-CATE targeting maximizes policy value at any budget and weakly dominates predictive (outcome-based) targeting, strictly when the two rankings disagree; (ii) a split-sample doubly robust evaluation of targeting quality is leakage-free (null-exact in finite samples), whereas the naive in-sample version is optimistically biased; and (iii) greedily targeting by CATE traces the optimal cost–benefit (Qini) frontier, with the deployment rule “recommend when τ^>0.” We validate the method on 17 randomized embedded experiments from the ASSISTments platform. Because the 40 held-out splits re-partition the same students, we do not treat them as independent replicates: we calibrate every headline comparison against a within-experiment permutation null and an experiment-clustered bootstrap. Under that calibrated inference, causal targeting outperforms predictive targeting for adaptive homework at every budget (permutation p≤0.005, the resolution floor of 200 replicates; Holm-corrected p≤0.040), while for learner choice the same contrast is directionally consistent but not statistically significant (permutation p=0.23–0.38; clustered p=0.42). The direction is stable in both families: no leave-one-experiment-out refit reverses its sign. Predictive targeting is nonetheless the one rule that is reliably worse than the alternatives, because it recommends the feature to high-performing students who benefit least—realized benefit falls monotonically across predicted-performance deciles (from +0.087 in the lowest to −0.030 in the highest). Against a fuller baseline suite, causal (CATE) targeting does not beat random, a simple risk-based rule (target low performers), or treating everyone. Indeed, the estimated benefit ranking is close to noise—split-half rank agreement is ρ≈0.002–0.008 and its calibration slope is 0.018, far below the ideal of 1—so the gain over predictive targeting comes from avoiding an actively harmful ordering rather than from recovering individual benefit. A fairness analysis shows why this matters: predictive targeting is regressive, concentrating feature access on high-ability students, whereas causal and risk-based targeting reverse that gradient in this corpus; no policy differentiates by neighborhood opportunity zone. Group-conditional policy values, however, are not individually distinguishable from zero once dependence across students and experiments is accounted for; what survives resampling is the allocation itself—predictive targeting directs 0.33 fewer of its recommendations to low-ability than to high-ability students (95% CI [−0.46,−0.01], experiment-clustered)—so we frame the fairness result as improved access, not established equity gains. The actionable finding is therefore narrow and specific: outcome-based targeting systematically mis-allocates learning-path features and should be replaced by some benefit-aware rule; whether that rule needs to be a learned CATE model, rather than a simple risk-based heuristic, is not established by this corpus.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

Efficient Offline Learning of Ranking Policies via Top-$k$ Policy Decomposition

Many recommender systems such as for e-commerce and news platforms aim to provide users with rankings they are likely to interact with. Off-Policy Learning (OPL) of ranking policies enables us to learn new ranking policies using only historical logged data. However, ranking settings make OPL remarkably challenging beca...

Ren Kishimoto, Koichi Tanaka, Haruka Kiyohara et al. · 0 citations
#machine learning Preprint Sep 2026

Stable Policy Learning

This paper proposes a method for policy learning called policy-vote bagging, which learns treatment decisions on many subsamples then averages their votes into treatment probabilities, which preserves expected welfare and improves expected utility for a risk-averse researcher.

Harvey Barnhard, Giacomo Opocher, Rahul Singh · 0 citations
Conference Open access 2026

Off-Policy Evaluation of Contextual Bandit Algorithms for Personalized News Recommendation Using the MIND Dataset

: News recommendation systems rely on user feedback to improve ranking decisions. Running online tests adds cost. It can also interrupt the user experience. Off-policy evaluation (OPE) gives a safer option for estimating new policies. The paper models news recommendation as a contextual bandit task. The paper construct...

Yi-Kai Zhao · 0 citations
#machine learning Preprint Sep 2026

Recommendation Ranking Off-Policy Evaluation under Ranking-Dependent Examination via Examination-Relevance Decomposition

Off-policy evaluation, which estimates evaluation policy performance from logged data, is key for recommender ranking policies. However, logged clicks cannot distinguish unexamined items from examined non-clicks, causing bias in existing estimators when the assumed examination structures fail. We propose two estimators...

Riki Okamura, Toshiharu Sugawara · 0 citations
Open access Sep 2026

Artificial Intelligence for Personalized Marketing in Digital Learning Platforms: Trade-Offs Across Neighborhood, Behaviorally Segmented, and Neural Recommender Models

Artificial intelligence (AI) recommender systems operationalize personalized marketing by transforming behavioral data into individualized choice architectures, yet greater model complexity or personalization granularity need not improve every service objective. This study compares population-level item-neighborhood re...

Nerantzoula Sevaslidou, Eugenia Papaioannou, K. Assimakopoulos et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.