Jul 2026· IEEE Robotics and Automation Letters· Vol 11, pp. 10736-10743· 0 citations· 34 references
Computer Science
TL;DR
This letter introduces a language-guided object removal that combines neural field resampling with multiview-consistent progressive inpainting, a direct NeRF weight editing method utilizing knowledge distillation, and the first benchmark (NEO-Dataset) for quantitatively evaluating NeRF scene editing methods suitable for robot manipulation.
Abstract
In this letter, we present NEO, a unified framework providing language-guided NeRF editing for robotic manipulation. Our letter introduces (i) a language-guided object removal that combines neural field resampling with multiview-consistent progressive inpainting, (ii) a direct NeRF weight editing method utilizing knowledge distillation, composing original and edited NeRFs via a teacher–student model, enabling coherent modeling of future scene states before a robot executes an action, and (iii) the first benchmark (NEO-Dataset) for quantitatively evaluating NeRF scene editing methods suitable for robot manipulation. We show that our approach outperforms state-of-the-art baselines in scene editing tasks, including object removal and pick-and-place robotic experiments, yielding visually coherent and geometrically consistent edits that reduce artifacts commonly introduced by prior methods. Finally, we showcase the capability of NEO for multi-stage robotic assembly tasks by preserving a persistent NeRF plus language-field representation after each edit, enabling iterative future-state scene representation prediction without requiring additional scene re-scanning.
Learning-based manipulation requires supervision that is both semantically meaningful and physically executable, but current data pipelines often provide only one of these properties. Human demonstrations capture intent but are costly to collect and constrained by the human-robot embodiment gap, while simulation can sc...
Yian Wang, Jun-Yi Cao, Xiao-Wen Qiu et al.· 0 citations
Embodied co-painting requires a robot to repeatedly update a shared physical canvas while human intent evolves over interaction. Existing reference-driven painters or reactive assistants are typically optimized for single-shot rendering or sketch completion, limiting their ability to sustain coherent multi-round collab...
Dan-Tong Qin, Yi-Ke Guo, Qin-Lin Liu et al.· 0 citations
General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual pre...
Hao-Ran Wen, Wen-Fu Wang, Kun-Song Shi et al.· 1 citation
Text-guided image editing aims to perform a desired edit while preserving source content unrelated to it. Pretrained rectified-flow models enable training-free editing of real images through modifications to their sampling trajectories. However, responses at locations unrelated to the desired edit can still accumulate...
Jing-Xuan Kang, Yin-Song Wang, Che Liu et al.· 0 citations
Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...
Trong-Tai Dam Vu, Vinh-Tiep Nguyen· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.