Stable Policy Learning
This paper proposes a method for policy learning called policy-vote bagging, which learns treatment decisions on many subsamples then averages their votes into treatment probabilities, which preserves expected welfare and improves expected utility for a risk-averse researcher.
Harvey Barnhard, Giacomo Opocher, Rahul Singh
· 0 citations