Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be...
Yi-Jia Fan, Zi-Qi Huang, Zhongang Cai et al.· 0 citations
We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully curated generation...
Hai-Wen Diao, Jia-Hao Wang, Chen-Jing Ding et al.· 0 citations
ReDeck is proposed, a step-level render-grounded refinement framework that decomposes slide revision into atomic edit actions and returns renderer-derived observations after each step, turning refinement into"one edit, one observation."
Mu-Zhao Tian, Ze-Zi Zeng, Yi-Fan Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.