Architectural floor plans remain a high-friction barrier to archive digitization and early design-model preparation because heterogeneous graphics encode spatial semantics and editable geometry together. We introduce SALI-FP, an evidence-gated multimodal pipeline that converts a plan into reviewable semantic maps, obje...
Hong-Xuan Chen, Wenda Wang, Jia-Chen Lu et al.· 0 citations
WorldClaw is presented, a fully agentic, coarse-to-fine framework for open-world 3D scene generation that produces large-scale scenes with coherent spatial organization, visually compelling local content, and editable instance-level assets while preserving a consistent global terrain structure.
Chunchao Guo, Jinpeng Li, Yang Li et al.· 5 citations
While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely, code-based visual generation via Coding Agents provides precise layout control...
Jun-Yan Ye, Wei Liu, Dongzhi Jiang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.