Skip to content
Preprint

GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning

Sep 2026 · 1 citation · ⚡ 1 influential · 115 references
Computer Science

TL;DR

GIF, an agentic Generation framework for Interactive and Functional object compositions, recast this problem as disentangled reconstruction followed by relative pose recovery, revealing diversity scaling in both simulation and real-world deployment.

Abstract

Robot manipulation foundation models require scalable evaluation and data generation across diverse scenarios, with simulation providing an environment for both. Automated scene generation offers a promising path, yet prior work has largely emphasized coarse-grained scene layouts rather than fine-grained functional object compositions. Motivated by this gap, we present GIF, an agentic Generation framework for Interactive and Functional object compositions. In this framework, we recast this problem as disentangled reconstruction followed by relative pose recovery. CoGen produces instance-disentangled meshes with coarse initial poses leveraging complementary strengths of 2D and 3D generative models. GPRM refines the relative pose under joint geometric and physical guidance, and a VLM verifier selects the candidate that best matches the structured specification. We further construct a benchmark spanning eight representative contact-geometry classes and compare with state-of-the-art generators; GIF improves both asset quality and relation matching, while reducing collision rate to below 1%. Finally, we synthesize data for policy learning, revealing diversity scaling in both simulation and real-world deployment.

View source

Similar papers

Preprint Aug 2026

iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits the usefulness of synthetic data for downstream training, where access...

Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel et al. · 1 citation
Preprint Aug 2026

Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction

RoomWright is presented, an agentic usage-driven framework for generating 3D scenes represented entirely as code for embodied interaction, providing interactive environments for embodied AI and policy learning.

Zijian Xiao, Zi-Peng Ye, Jin-Kun Hao et al. · 1 citation
Preprint Sep 2026

MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence

General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual pre...

Hao-Ran Wen, Wen-Fu Wang, Kun-Song Shi et al. · 1 citation
Preprint Sep 2026

IM-ENGINE: Image Editing for Embodied Data Generation

Learning-based manipulation requires supervision that is both semantically meaningful and physically executable, but current data pipelines often provide only one of these properties. Human demonstrations capture intent but are costly to collect and constrained by the human-robot embodiment gap, while simulation can sc...

Yian Wang, Jun-Yi Cao, Xiao-Wen Qiu et al. · 0 citations
#machine learning Preprint Sep 2026

Learning Expressive and Compositional Motion Representation via Spectral Skills

Robotic foundation models offer a promising path toward general-purpose humanoid robot control, often through hierarchical architectures. However, their effectiveness depends on the command interface between the planner and the controller, which must support accurate execution while remaining easy to predict, and ideal...

Fei-Yang Wu, Chen-Xiao Gao, Chen Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PhysMAS: Physics-Grounded Multi-Agent Synthesis of Compositional 4D Gaussians

Efficient, fully automatic, and physically plausible 4D Gaussian synthesis is an important goal for dynamic scene generation. Recent physics-based methods couple 3D Gaussians with the Material Point Method (MPM) to generate physically driven motion, but extending this paradigm to heterogeneous multi-part objects and in...

Jiang Qin, Chun-Ji Lv, Yang-Guang Wei et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.