VoT: Vision-of-Thought for Unified Multimodal Representation Alignment
Current text-to-image systems typically employ a"text encoder plus diffusion decoder"paradigm, in which text semantics directly modulate continuous latent noise. Despite their success, these methods lack an explicit, interpretable intermediate representation that effectively bridges high-level linguistic semantics and...