LIFD: Anchored Diffusion for 3D-Aware Scene Memory in Robotic Manipulation
This work introduces \lifd{} (Look, Imagine, Focus, and Do), a framework for persistent, 3D-aware scene memory that learns scene tokens through multi-view agreement, then completes them from a single RGB view and recurrent memory using rectified flow.