Preprint
Jul 2026
Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL
Three RLxF principles are argued to apply equally to model-based control and to verifier-based or RLHF approaches in LLM alignment and include ground risk in world outcomes, validate proxies before deployment, and substitute outcome-trained feedback models when direct world signals are unavailable.
Zhaohui Wang
· 2 citations