Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these metrics evaluate perceptual dimensions sep...
Yin-Ming Huang, Shu-Yuan Tu, Xi Yan et al.· 4 citations
Achieving virtual agents capable of automating tasks across diverse digital environments remains a pivotal challenge in Embodied AI. While Multimodal Large Language Models (MLLMs) offer enhanced visual perception and reasoning, their agentic deployment faces three challenges: costly data annotation, imprecise action-in...
Jia-He Ying, Wendong Bu, Kaihang Pan et al.· 0 citations
This work states that existing visual generative models are not yet ready for RL due to the following two fundamental drawbacks that undermine the foundations of RL.
Bohan Wang, Min Zhou, Zhongqi Yue et al.· Neural Information Processin...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.