Learning to Allocate Incentives for Incentivized Advertising via Offline Model-Based Reinforcement Learning
An offline model-based RL framework for cost-controllable sequential incentive allocation is developed and an independent counterfactual scorer evaluates each learned policy on held-out logs, enabling pre-launch selection without costly online exposure.
Zi-Lin Zhao, Han Yang, Tianpei Yang et al.
· 0 citations