Skip to content

Author

Hai D. Pham

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Jul 2026

Combining Policy Gradients with Quality-Diversity in Cooperative Multi-Agent Reinforcement Learning

Quality-Diversity (QD) methods combined with policy gradients have shown strong performance in single-agent reinforcement learning, but extending them to multi-agent settings introduces challenges from partial observability and agent interactions. We propose MAPGA-ME, a multi-agent extension of PGA-MAP-Elites that integrates policy gradient updates into MAP-Elites for cooperative control. Our results show that directly transferring policy gradient mechanisms from single-agent QD does not consistently improve performance in multi-agent environments. In particular, a design choice effective in single-agent settings becomes less suitable under decentralized, partially observable conditions. Across multiple configurations, we identify key factors affecting the effectiveness of policy gradient-based QD in multi-agent learning, providing practical guidance for adapting these methods.

Hai D. Pham, Ngoc Hoang Luong · 0 citations