Complex 3D spatial text to image generation requires models to convert natural language into stable visual geometry, not merely semantic appearance. Existing prompt-driven or layout-conditioned methods improve controllability, but often lack an optimizable and verifiable spatial intermediary before visual sampling. As...
Ziyun Qian, Zi-Zhi Chen, Yi-Zhou Liu et al.· 2 citations
Generative models have advanced image-conditioned 3D content creation, yet generating controllable and executable 3D scenes from a single image remains challenging. Existing 3D generative approaches can synthesize visually plausible objects and scenes, but their spatial layout estimation is coupled with specific asset...
Ding-Kang Yang, Yi-Zhou Liu, Wen-Dong Cheng et al.· 0 citations
Omni-modal models have expanded multimodal interaction across vision, audio, speech, and language. However, their training is predominantly organized around semantic descriptions and general-purpose objectives, leaving physical attributes, interaction states, and causal mechanisms only partially specified. This gap is...
Yi-Zhou Liu, Jing-Hang Han, Kai Qiu et al.· 0 citations
TRACER is proposed, which formulates compression as a sequential per-tool decision problem, and demonstrates the value of consequence-aware, per-tool context retention for improving the efficiency of long-horizon language agents.
Zihao Lin, Ye Wu, Mengning Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.