Skip to content
Preprint

Better Slots, Better Worlds: Representation Quality&Robustness in Object-Centric World Models

Aug 2026 · 0 citations · 32 references
Computer Science

TL;DR

Under unseen distribution shifts, the OCWM with well-bound slots is more robust overall than the end-to-end trained scene-centric LeWM, while DINO-WM, built on similar frozen pretrained features, remains competitive -- pointing to pretrained features as a key contributor to robustness.

Abstract

Learning world models from offline trajectories enables agents to accomplish different tasks through planning. Object-centric (OC) representations, which decompose a scene into a set of slots that bind to its objects, have been proposed as an inductive bias for world models that are more sample-efficient and generalize better. Yet prior object-centric world models (OCWMs) take the slot encoder as given and evaluate only in-distribution, leaving open whether the object-centric bias actually delivers for planning and what within the OCWM drives it. We conduct a controlled study of OCWMs for visual model-predictive control along two axes: object-centric representation quality and generalization under distribution shift relative to scene-centric models. We find that (i) planning success correlates positively with unsupervised slot-quality metrics (FG-ARI, mBO), though the gains saturate at high slot quality; (ii) with well-bound slots, the auxiliary proprioception inputs and masking inductive bias that prior methods relied on become unnecessary; and (iii) under unseen distribution shifts, the OCWM with well-bound slots is more robust overall than the end-to-end trained scene-centric LeWM, while DINO-WM, built on similar frozen pretrained features, remains competitive -- pointing to pretrained features as a key contributor to robustness.

View source

Similar papers

Open access Sep 2026

Vision-Centric World Models for Embodied Robots: Representations, Predictive Interfaces, and Evaluation

Embodied robots need more than a description of the current image: they must estimate how the scene may change under motion, contact, and partial observation. Vision-centric world models provide this predictive layer, but they expose it through different state interfaces. This paper organizes the literature into four f...

Yi-Ning Li · 0 citations
Preprint Sep 2026

Object-Centric Conditioning for Visuomotor Flow Matching

Robot visuomotor policies are commonly formulated as autoregressive, diffusion-based, or more recently, flow matching models. Among them, Action-to-Action (A2A) flow matching improves inference efficiency by initializing generation from historical action priors rather than stochastic noise. However, stale historical mo...

Ji-Jie Li, Xu Yang, Jun-Hong Zou et al. · 0 citations
Preprint Aug 2026

presto: Efficient, Training-free, and Open-world Object Placement via Imaginary Search

This work reformulates open-world object placement as a heuristic search task guided by reasoning from a Multimodal Large Language Model (MLLM), and introduces Presto, a zero-shot, training-free framework that operates within an imaginary action space to iteratively refine object position and scale.

Weixuan Ding, Shang Liu, Han-Yu Pei et al. · 0 citations
Preprint Sep 2026

The Planning Limits of Latent World Models

World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet existing studies mainly demonstrate what these models can accomplish, leaving unclear when their predictions remain useful for planning and where they fail. We study this ques...

A. Alrasheed, Basim Azam, Naveed Akhtar · 0 citations
#artificial intelligence Preprint Sep 2026

DeepJEPA: Scaling World Models from Within

World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition. We show that making every transition uniformly deeper wastes computation and can degrade planning because useful refinement is concentrate...

Zi-Jian Jin, Yun-Bei Zhang, Yuan-Zhe Liu et al. · 0 citations
May 2025

Object Concepts Emerge from Motion

Object concepts play a foundational role in human visual cognition, enabling perception, memory, and interaction in the physical world. Inspired by findings in developmental neuroscience - where infants are shown to acquire object understanding through observation of motion - we propose a biologically inspired framewor...

Hao Liang, Xiao-Hui Wang, Zhi-Chao Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.