Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choi...
Hong-Yang Du, Yun-Fei Xie, Jun-Jie Ye et al.· 0 citations
Dynamic Gaussian Splatting provides an explicit representation of evolving 3D scenes, but existing approaches are primarily optimized for reconstruction, future-state generation, or rendering rather than for learning reusable predictive dynamics. We propose 4DGS-JEPA, a Gaussian-native joint-embedding predictive archit...
Joint-embedding predictive architectures learn abstract states by predicting target embeddings from context embeddings from context embeddings, but their transition models are typically opaque neural maps, so SJEPA is introduced, a reconstruction-free JEPA framework that learns predictive representations whose induced...
The introduction of Markov-Chain JEPA (MCJEPA), which replaces the latent predictor by a learned transition matrix, and Markov-Chain JEPA (MCJEPA), which replaces the latent predictor by a learned transition matrix, give temporal JEPA a principled state-space interpretation.
Yong-Chao Huang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.