Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference. In this work, we develop a simple and effective post-training recipe for open-domain text-to-image gen...
Yuan-Hao Ban, I-Hung Hsu, A. Angelopoulos et al.· 0 citations
One-Forcing is proposed, a simple yet effective approach that augments the DMD objective with an auxiliary GAN loss for high-quality and efficient one-step video generation, and finds that framewise autoregression stabilizes adversarial training, enabling higher-quality generation with substantially fewer training iter...
Jia-Qi Feng, Justin Cui, Yuan-Hao Ban et al.· arXiv.org· 13 citations· ⚡3
SlackDrive is proposed, a pre-inference compute allocator that reuses realized latency to select the compute budget of each control step before model execution, complementing existing profiling and resource scheduling while preserving the driving backbone and its compute actuator.
Xiao-Huan Pei, Heng-Guang Zhou, Yuan-Hao Ban et al.· 0 citations
Improvements are shown that ReCAST yields improvements that generalize beyond the training rewards and support its core principle: assigning each reward greater weight at the denoising timesteps where its feedback is most informative.
Yi-Hang Chen, Yuan-Hao Ban, Kuei-Chun Kao et al.· 0 citations
This work replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards, and adds a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts.
Yuan-Hao Ban, Jia-Qi Feng, Heng-Guang Zhou et al.· 0 citations
Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models. Existing faithfulness benchmarks, however, rely on simple atomic instructions, on which top-tier systems already achieve near-perfect scores. As T2I models enter cre...
Yuanhao Ban, Tong Xie, Sohyun An et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.