Action-conditioned video world models predict future robot interactions from multiple cameras, yet their outputs remain disparate video collections rather than a shared metric scene queryable across viewpoints and time. While existing 4D reconstruction methods offer a path to spatialize these predictions, independently...
Jin Hyun Kim, Min Young Kim, Soohwan Song et al.· 0 citations
3D scene graphs connect spatial perception with high-level language reasoning, but unverified false-positive groundings can impair downstream subtask generation. We propose a feed-forward, spatially constrained framework that grounds task-relevant objects and generates position-aware subtasks without an iterative task–...
Deep learning-based Multi-View Stereo (MVS) has advanced significantly but often generalizes poorly to unseen scenes, particularly in occluded areas or regions with limited view overlap. To mitigate this, recent approaches integrate Depth Foundation Models (DFMs) into MVS pipelines to provide monocular depth priors. Ho...
Byeonggwon Lee, Sanggil Lee, Siwoo Lee et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.