Vision-language pre-training has reshaped image clustering, giving rise to language-assisted image clustering (LaIC), which leverages textual semantics to complement visual representations. Despite the rapid proliferation of LaIC methods, it remains unclear how much LaIC has actually advanced image clustering, as exist...
We introduce SolarWM, a fully open foundation for building interactive video world models from data preparation through long-horizon inference. Training across heterogeneous data sources and video backbones is challenging: datasets differ in temporal scale, camera geometry, visual quality, motion, and captioning styles...
Jun-Chao Huang, Gui-An Fang, Sheng-Ju Qian et al.· 1 citation
VideoRAE is introduced, a representation autoencoder that converts features from a frozen video foundation model into compact, reconstruction-capable latents for video generation, establishing frozen video foundation representations as compact, versatile, and generation-friendly video latents.
Zhihao Xie, Junfeng Wu, Xinting Hu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.