A learned simulator can reproduce its training conditions accurately yet fail in two distinct ways once those conditions change. Over long rollouts, small errors accumulate until the trajectory drifts away from physically plausible behavior; under an intervention on a physical parameter, the model may continue to follo...
Yu-Feng Wang, Parivesh Priye, Lu Wei et al.· 0 citations
Bayesian quantum tomography requires efficient inference while preserving a posterior fixed by the prior and Born likelihood. Learned transport provides fast amortized samples, but reward tuning can reshape the generated distribution rather than improve exploration of this fixed target. We introduce GRPO-QPS, a target-...
Yu-Feng Wang, Parivesh Priye, Lu Wei et al.· 0 citations
The diagnosis prescribes the fix: keep the goal out of the dynamics and supervise the \emph{read} path, recovering genuine, instruction-independent grounding, and the detection protocol and remedy apply to any goal-conditioned world model whose instruction names the scored quantity.
Yufeng Wang, Lu Wei, Haibin Ling· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.