Skip to content
Conference Open access

PRIME: A Decoupled Multi-agent Actor-Critic for Multi-view Clustering

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · pp. 4295-4303 · 0 citations · 27 references

TL;DR

A decoupled multi-agent actor-critic (PRIME) is proposed via defining multi-view clustering as a partially observable Markov game, which establishes a dynamic trial-and-error assignment process of samples in a decentralized perception with centralized learning framework to adaptively learn the optimal clustering strategy.

Abstract

Deep multi-view clustering draws plentiful attention in various domains, owing to remarkable performance in learning patterns from complementary information of multi-view data. However, previous methods encounter two challenges. They utilize a single pre-defined clustering strategy to perceive diverse structures from data in multiple views subject to heterogeneous distributions, failing to fully capture intricate complementary structures. They leverage either the feature fusion or the result fusion in clustering, which cannot fully integrate complementary information. Therefore, a decoupled multi-agent actor-critic (PRIME) is proposed via defining multi-view clustering as a partially observable Markov game, which establishes a dynamic trial-and-error assignment process of samples in a decentralized perception with centralized learning framework to adaptively learn the optimal clustering strategy. In PRIME, an actor leverages the policy gradient paradigm to independently implement a Markov decision process of data partition in a view, which fully explores data structures to tailor a clustering sub-policy for a view. Meanwhile, a critic utilizes the value function paradigm to centrally guide the Markov game among actors in different views, which constructs both feature and result fusion to progressively enhance complementary knowledge integration for robust clustering results. Extensive experiments on 6 benchmark datasets verify the superiority of PRIME against 10 methods.

Read PDF

Similar papers

#machine learning Preprint Oct 2026

Test-time Multi-agent Coordination by Decomposed Value Gradient Flow

Offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data, but cannot distinguish high-value regions, while value-optimized policies exploit the learned Q-function but collapse the multi-modal into a single dominant...

Dong-Su Lee, Hao-Ran Xu, Amy Zhang · 0 citations
Preprint Aug 2026

One Model, Many Minds: Unlocking Multi-Agent Synergy in a Single Agent via Mixture of Roles

The proposed Mixture of Roles (MoRe), which adaptively composes multiple specializations into a single steering vector for single-turn inference, enables multi-perspective specialization in a single-agent, single-turn inference process.

Zhichen Zeng, Hui-Yuan Chen, Jingru Cheng et al. · 2 citations
#machine learning Preprint Sep 2026

HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning

Cooperative multi-agent reinforcement learning under partial observability and shared rewards requires assigning team outcomes to individual agents and high-order coalitions. A MAPPO-style critic compresses joint behavior into one global value, while critics that dynamically reconstruct the grouping topology change the...

Xing-Long Luo, Yu-Ding Zhang, Yu-Heng Kuang et al. · 0 citations
Open access Aug 2026

AGTA: Topology-Aware Sequential Decision-Making in Multi-Agent Reinforcement Learning

Action Generation with Topology Awareness (AGTA), a topology-aware sequential decision-making framework in MARL that integrates inter-agent correlation modeling with topology-guided decision-order optimization, and outperforms the state-of-the-art counterparts.

Kun Hu, Shanghua Wen, Wen-Di Wu et al. · 0 citations
#machine learning Preprint Sep 2026

ICMAPE: In-Context Multiagent Pure Exploration

ICMAPE converts the fixed-confidence identification objective into a reward derived from inference confidence, so that standard reinforcement learning machinery can be applied to decentralized pure exploration.

Xin-Yi Hu, Alessio Russo, Aldo Pacchiano · 0 citations
Preprint Aug 2026

Test-Time Collaborative Classification over Multi-Agent Networks

The increasing heterogeneity of multi-agent systems poses significant challenges for jointly training a global model across agents. At the same time, cooperative inference between agents has long been recognized as a powerful mechanism for distributed decision making over networks. Motivated by these observations, we p...

Ping Hu, Mert Kayaalp, Ali H. Sayed · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.