Recent diffusion and autoregressive models have substantially improved text-to-video generation, yet producing coherent long-form story videos with consistent characters and aligned audio remains challenging. Existing methods often suffer from narrative drift, unstable character identity, weak cross-scene continuity, a...
Ting Huang, Biao Wu, Rong-Hao Chen et al.· 0 citations
While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shap...
Shuo Zhang, Yifan Zhou, Han-Yu Wang et al.· 0 citations
This work presents \textsc{SemaPLC}, a project-grounded and verification-gated agent harness assembled from conventional tools but governed by a strict completion rule, which raises the mean at every layer and most sharply at runtime.
Yan-Lun Tu, Hua-Can Wang, Zi-Yue Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.