Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human kn...
Jun-Wei Zhou, Zhen Sun, Binyu Li et al.· 2 citations
Diagnostics reveal that RL on PLMs is governed by two reward properties: verifiability, whether the reward is a fixed environment or a learned surrogate vulnerable to distribution shift, and coverage, the fraction of sequence space giving an informative gradient.
Hanqun Cao, Hongrui Zhang, Junde Xu et al.· Proceedings of the 32nd ACM...· 0 citations
This work introduces Emotion Statement Judgement (ESJ), a statement-verification formulation that preserves the expressiveness of the input space while constraining outputs to discriminative judgements, and builds EmObserver, an emotion-oriented MLLM optimized on ESJ through an elaborate multi-stage recipe.
Daiqing Wu, Dong-Bao Yang, Jiashu Yao et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.