Process reward models (PRMs) have demonstrated notable effectiveness in test-time scaling and reinforcement learning by providing fine-grained signals for evaluating intermediate reasoning states, but their training relies heavily on costly process annotations. A natural way to alleviate this dependence is to complemen...
Kai Gan, Zi-Hao Zhou, Bo-Tao Ye et al.· 0 citations
Recent advances in multimodal large reasoning models (MLRMs) have demonstrated impressive capabilities on complex multimodal tasks, yet their reliance on long Chain-of-Thoughts (CoTs) often leads to redundant reasoning and high computational cost. Existing chain-based distillation and refinement approaches alleviate re...
Yizhi Wang, Li-Nan Yue, Deng-Bao Wang et al.· Proceedings of the 32nd ACM...· 0 citations
This work introduces MRCL, a Multimodal Reasoning Continual Learning benchmark, and proposes Continual Policy Optimization (CPO), a replay-free framework grounded in a prior-task behavioral KL objective that consistently reduces forgetting while preserving, and in some cases improving, pretrained model capabilities.
This work proposes ComRank, a ranking loss framework for MLCLL, which encourages complementary labels to be ranked lower than non-complementary ones, thereby modeling pairwise label relationships and ensures Bayes consistency under both uniform and biased cases.
Jin Zhu, Yi Gao, Miao Xu et al.· Neural Information Processin...· 0 citations
This work introduces a collective evidence-threshold backdoor paradigm for MAS and Boundary-Conditioned Backdoor Injection, which constructs counterfactual boundary pairs to separate benign behavior before the threshold from the adversarial objective after it, and learns latent progression aligned with evidence.
Jiahao Xiao, Lei Feng, Min-Ling Zhang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.