Modern stereo matching models fail when disparity crosses zero, with end-point error (EPE) rising by 4.6-37$\times$. Yet stereoscopic content, from cinema 3D to VR, routinely contains objects behind the zero-disparity plane (ZDP), corresponding to negative disparities. The blind spot cascades through datasets, architec...
Jian Shi, Xin-Ge Yang, Chao-Yang Wang et al.· 0 citations
We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D. Given a monocular video and an event-triggered prompt of the form"when X happens, do Y,"EgoPlay infers whether and w...
Jinjie Mai, G. Qian, W. Menapace et al.· arXiv.org· 0 citations
PhysStream is an autoregressive model for physics-grounded image-to-video synthesis that incorporates structured scene memory and supports fine-grained motion control via sparse velocity-increment signals that encode physical quantities, letting the model learn the underlying dynamics.
Chuhao Chen, Peter Wonka, Chao-Yang Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.