Spatial representation learning for autonomous driving aims to map raw visual signals into structured 3D scene representations, where object-centric bounding boxes and rendering-oriented 3D primitives (\eg, 3D Gaussians) serve as two distinct yet highly complementary levels for scene understanding. Existing methods typ...
Li-Heng Chen, Haokai Pang, Cheng Su et al.· 0 citations
This work proposes IQA-T1, a tool-based visual evidence reasoning framework that augments MLLM reasoning with explicit perceptual observations and constructs Q-Tool, a dataset containing 11k multimodal reasoning chains grounded in tool-generated evidence.
Remember-R1 is proposed, a reinforcement learning framework that mitigates long-context visual forgetting by applying process-level supervision directly on the original reasoning trajectory, demonstrating its effectiveness in mitigating long-context visual forgetting.
Jianmin Chen, Jiaqi Tang, Wei Wei et al.· 0 citations
Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-only conditioning and creation-only training do not explicitly supervise the local structure needed for precise, temporally consistent editi...
Zhe-Fan Rao, Bin-Yi Zou, Xuanhua He et al.· 0 citations
MSEditor is proposed, the first framework designed specifically for consistent multi-shot video editing, which significantly outperforms existing methods on the authors' curated multi-shot video editing benchmark in terms of identity preservation, temporal stability, and overall visual quality.
Kunyu Feng, Yue Ma, Bing-Yuan Wang et al.· 1 citation
This work organizes the embodied data ecosystem as a pyramidspanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelit...
Yifan Ye, Yankai Fu, Ya-hui Lv et al.· arXiv.org· 4 citations
LeapBot-WA establishes a novel Predictive-Latent paradigm for WAMs by operationalizing the Joint-Embedding Predictive Architecture (JEPA) as a World-Anchor and introduces the Isotropic Semantic Autoencoder (ISAE), which reshapes the anchor's latent space into a diffusion-friendly manifold to prevent off-manifold drift.
Pei Liu, Nan Zheng, Lang Zhang et al.· arXiv.org· 0 citations
This work investigates LLM-based forecasting agents, meaning systems in which a language model contributes to a scored prediction about a future or currently unobserved target, and organizes architectures into three groups.
Xiao-Gang Xu, Jiaqi Tang, Jianmin Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.