Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics...
Abhinav Sharma, S. Navuluru, Wang Wei et al.· 0 citations
Air-ground transportation research increasingly relies on co-simulation, yet constructing scenarios remains labor-intensive and difficult to validate. More importantly, a generated scenario may execute successfully while failing to realize the spatial, temporal, communication, or behavioral relationships requested by t...
Visko Orbis 1.0 achieves the best DOVER aesthetic and technical scores and the best VideoAlign visual and motion quality, and leads three physical-plausibility protocols (VideoPhy-2, Physics-IQ, and VBench-2.0 Physics); in long-form Arena comparisons, it obtains the highest overall-preference and temporal-stability rat...
Xiang-Bo Gao, Siyuan Yang, Ping He et al.· arXiv.org· 5 citations
This survey delivers a comprehensive and critical synthesis of the emerging role of GenAI across the autonomous driving stack, delving into the frontier applications of GenAI in image, LiDAR, trajectory, occupancy, and video generation, as well as LLM-guided reasoning and decision-making.
Yu-Ping Wang, Shuo Xing, C. Cui et al.· ACM Computing Surveys· 60 citations· ⚡2
Agentic Chain-of-Thought Steering (ACTS), which formulates reasoning steering as a Markov decision process where a controller agent adaptively steers a frozen reasoner during inference, and enables budget-aware strategy control for efficient reasoning while preserving the reasoner's generation continuity.