OmniVAE is presented, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations that translates into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation.
Jun Zhan, Chenchen Yang, Y. Gong et al.· 0 citations
ContextWeave is introduced, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams and motivates memory systems that optimize not only retrieval relevance but also reliable use during execution.
AgentHPOBench, a sequential benchmark comprising 30 executable machine learning tasks across seven research categories, shows that current agents exhibit measurable experimental optimization ability across domains, but still face clear limitations in sustained iterative refinement, complex log diagnosis, and consistent progress toward reported reference performance.
Tianyu Huai, Tingshuo Fan, Xinchi Chen et al.· 0 citations