Preprint
Aug 2026
Learning to Control Coupled-Dynamics Environments with Joint Markov Decision Processes
A nonparametric distributional Bellman optimality operator for JMDPs is defined, and it is proved that when the induced marginal MDP has a unique optimal policy, its iterates converge in Wasserstein distance to the optimal joint return law.
Ege C. Kaya, Aliasghar Pourghani, Mahsa Ghasemi et al.
· 0 citations