Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts. We trace these failures to an early coordination bottleneck: before denoising begins, prompt-conditioned attention may allocate different concepts to strongly overlapping spati...
Ning Zhu, Anchi Chen, Mengfei Zhao et al.· 1 citation
Few-step distillation reduces the inference cost of image-to-video generation, but directly reusing LoRA adapters trained for long denoising trajectories can weaken their intended effects and degrade video quality. We observe that adapters with similar measured static parameter geometry can behave differently under the...
Shi-Hong Li, Jun-Tao Xu, Cao Jin et al.· 0 citations
Cold-Start Active Learning (CSAL) aims to select a valuable subset from an unlabeled pool without any prior knowledge or human assistance. Existing methods take diverse routes based on typicality, coverage, or diversity. Each rests on its own inductive bias and therefore performs well on some tasks yet poorly on others...
Ning Zhu, Xiaochuan Ma, Jun-Tao Xu et al.· 0 citations
Pairwise preference labels rank complete images, yet Diffusion-DPO applies their effect over many spatial and denoising-time coordinates. For attention-based, noise-prediction latent diffusion, ToPO (Token-Oriented Preference Optimization) constructs a per-minibatch, detached, separable spatial-temporal route from bran...
Jun-Tao Xu, Shi-Hong Li, H. Au et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.