In neural decoding research, reconstructing natural images from fMRI signals poses a captivating yet challenging problem. Conventional approaches use basic linear mapping functions to project fMRI signals into a prior latent space (e.g., image and text embeddings) and subsequently utilize a pre-trained image generation...
Bich-Nga Pham, Trong-Tai Dam Vu, Anh-Khoa Nguyen Vu et al.· International Conference on...· 0 citations
Compositional text-to-image generation requires faithful depiction of multiple objects with distinct visual attributes. While recent inference-time optimization methods have substantially improved attribute binding, entity neglect - where specified objects are absent from the generated image - remains an unresolved fai...
Trong-Tai Dam Vu, Vinh-Tiep Nguyen· International Conference on...· 0 citations
Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...
Trong-Tai Dam Vu, Vinh-Tiep Nguyen· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.