Preprint
Aug 2026
Submodular Policy Learning for Distributed Task Allocation in Open Multi-Agent Systems
It is proved that the marginal gains of the stage utility provide an unbiased estimator of the gradient of the PME and that maximizing the PME over action distributions is equivalent to maximizing the stage utilities over agent actions, which are critical to devise principled policy gradient.
Jing Liu, Luca Ballotta, Yangyang Yang et al.
· 0 citations