Skip to content
Preprint

presto: Efficient, Training-free, and Open-world Object Placement via Imaginary Search

Aug 2026 · 0 citations · 59 references
Computer Science

TL;DR

This work reformulates open-world object placement as a heuristic search task guided by reasoning from a Multimodal Large Language Model (MLLM), and introduces Presto, a zero-shot, training-free framework that operates within an imaginary action space to iteratively refine object position and scale.

Abstract

Object placement is critical in image composition, requiring spatially and semantically coherent positioning of objects within diverse scenes. Existing approaches typically rely on hand-crafted rules or supervised learning on limited datasets, which restricts their generalization and interpretability, especially in open-world scenarios involving novel objects and scenes. In this work, we reformulate open-world object placement as a heuristic search task guided by reasoning from a Multimodal Large Language Model (MLLM). We introduce \textsf{presto}, a zero-shot, training-free framework that operates within an imaginary action space to iteratively refine object position and scale. Our coarse-to-fine search strategy ensures fast convergence, and we evaluate two decision-making variants: Metric-guided Selection and MLLM-as-a-judge. Experiments across multiple benchmarks show that \textsf{presto}~achieves state-of-the-art performance, particularly in previously unseen, open-world settings. Human studies further reveal that the MLLM-as-a-judge variant produces more perceptually coherent placements than metric-driven approaches, highlighting a gap between standard evaluation metrics and human visual judgment.

View source

Similar papers

Preprint Aug 2026

WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes

Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is absent, large vision-language models usually address this task. We ask whether the same capability is available far more cheaply, from representations already learned by a...

K. Gupta, Ahmed Rafi Hasan, Md. Mahfuzur Rahman et al. · 0 citations
Preprint Aug 2026

Better Slots, Better Worlds: Representation Quality&Robustness in Object-Centric World Models

Under unseen distribution shifts, the OCWM with well-bound slots is more robust overall than the end-to-end trained scene-centric LeWM, while DINO-WM, built on similar frozen pretrained features, remains competitive -- pointing to pretrained features as a key contributor to robustness.

Shukrullo Nazirjonov, Sai Prasanna, Anna Manasyan et al. · 0 citations
#small language model Preprint Aug 2026

OVIP-SG: Open-Vocabulary Instance-Preserving Scene Graphs for Mapping and Retrieval of Small, Fine-Grained Objects

OVIP-SG is presented, a unified framework for instance-preserving semantic mapping, functional scene partitioning, and language-guided small, fine-grained object retrieval that outperforms ConceptGraphs under a unified evaluation protocol on Replica.

Tianjing Hao, Hai-Yu Lan, Ang Li et al. · 0 citations
Conference 2026

Zero-Shot Multi-Reference Personalization via MLLMs-Guided Layout Planning

This work proposes RIG (Regional Image-prompt Generation), a novel training-free framework for multi-reference personalized generation that significantly outperforms state-of-the-art adapter methods in terms of both personalization fidelity and text-layout alignment.

Junhao Feng · 0 citations
Conference Aug 2026

Open-Vocabulary Visual Relationship Detection Via Vision-Language Models And Attention Mechanisms

Traditional Visual Relationship Detection (VRD) systems are strictly bound to predefined, closed-set vocabularies, limiting their practical application in real-world environments where object interactions are highly diverse. Moving towards an open-vocabulary setting (OV-VRD) provides more flexibility but introduces the...

N. Nguyen, Trung Tran · 0 citations
Open access Sep 2026

Open-Vocabulary Instance Segmentation for Scene Understanding in Mobile Robots

An open-vocabulary semantic mapping pipeline is presented that integrates TALOS (TAgging–LOcation–Segmentation–Segmentation) with the probabilistic, instance-aware Voxeland framework and shows a better balance between map completeness, geometric clarity, semantic coherence, and instance separation.

Macoris Decena-Gimenez, Pepe Ojeda, J. Ruiz-Sarmiento et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.