Skip to content
Preprint

WorldClaw: Agentic 3D Open-World Generation at Scale

Aug 2026 · 5 citations · 110 references
Computer Science

TL;DR

WorldClaw is presented, a fully agentic, coarse-to-fine framework for open-world 3D scene generation that produces large-scale scenes with coherent spatial organization, visually compelling local content, and editable instance-level assets while preserving a consistent global terrain structure.

Abstract

Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstream editing and reuse. We present WorldClaw, a fully agentic, coarse-to-fine framework for open-world 3D scene generation. Planning agents translate a text prompt into a structured specification of regions, terrain, assets, materials, and spatial relations. WorldClaw then builds a globally coherent terrain foundation from semantic layouts, reusable assets, generative or procedural materials, and a region-aware height field. For detail-demanding regions, it generates terrain-conditioned compositions, reconstructs editable textured meshes, and recovers their placement on the terrain; render-based agents further refine terrain, objects, appearance, and contacts. Across diverse open-world prompts, WorldClaw produces large-scale scenes with coherent spatial organization, visually compelling local content, and editable instance-level assets while preserving a consistent global terrain structure.

View source

Similar papers

Preprint Sep 2026

WorldWeave: Growing Persistent Geometric Worlds for Video Generation

Despite rapid progress, world models still lack explicit, persistent structural memory, making it difficult to preserve consistent world structure during continual scene expansion and cross-view revisits. To address this limitation, we present WorldWeave, a world generation framework that decouples world-state maintena...

Yi-Fan Huang, Li-Fan Jiang, Qing-Yue Hao et al. · 0 citations
Preprint Aug 2026

4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied Simulation

Embodied agents need environments that are visually diverse, physically interactive, and changing over time. Procedural simulators can generate large interactive scene collections, and recent 4D generators produce compelling visual dynamics. Combining these properties in one environment, however, still demands extensiv...

Ze-Hao Qi, Hao-Chen Luo, Jia-Wang Bian et al. · 0 citations
Preprint Sep 2026

Fysiverse-3D-SimReady Technical Report: Agentic Physical Simulation for Pragmatic 3D World Reconstruction

Agentic recognition requires visual perception to move beyond static scene understanding and produce structured scene representations that support the perception--reasoning--action loop. Existing single-image 3D generation methods, however, mainly produce visually plausible object assets rather than simulation-ready sc...

Lin-Tao Wang, Ming-Yang Sun, Yang Liu et al. · 0 citations
Preprint Sep 2026

HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis

Compositional 3D scene reconstruction has recently been explored from two directions: agentic reasoning that provides semantic understanding of spatial relationships but lacks precise alignment with input images; and visual geometry foundation models that predict dense point maps from input images but the reconstructio...

Shuoyao Sun, Chen Wang, En-Xin Song et al. · 0 citations
Preprint Sep 2026

WorldSculpt: Generating Compositional Worlds from Grounded Videos

We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics....

Mu-Yao Niu, Ji-Xuan He, Rui-Han Yu et al. · 0 citations
Preprint Sep 2026

Fysiverse-3D-Vision Technical Report: Generating Executable 3D Worlds from Images through Unified Spatial Reasoning

Generative models have advanced image-conditioned 3D content creation, yet generating controllable and executable 3D scenes from a single image remains challenging. Existing 3D generative approaches can synthesize visually plausible objects and scenes, but their spatial layout estimation is coupled with specific asset...

Ding-Kang Yang, Yi-Zhou Liu, Wen-Dong Cheng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.