This work reformulates open-world object placement as a heuristic search task guided by reasoning from a Multimodal Large Language Model (MLLM), and introduces Presto, a zero-shot, training-free framework that operates within an imaginary action space to iteratively refine object position and scale.
Abstract
Object placement is critical in image composition, requiring spatially and semantically coherent positioning of objects within diverse scenes. Existing approaches typically rely on hand-crafted rules or supervised learning on limited datasets, which restricts their generalization and interpretability, especially in open-world scenarios involving novel objects and scenes. In this work, we reformulate open-world object placement as a heuristic search task guided by reasoning from a Multimodal Large Language Model (MLLM). We introduce \textsf{presto}, a zero-shot, training-free framework that operates within an imaginary action space to iteratively refine object position and scale. Our coarse-to-fine search strategy ensures fast convergence, and we evaluate two decision-making variants: Metric-guided Selection and MLLM-as-a-judge. Experiments across multiple benchmarks show that \textsf{presto}~achieves state-of-the-art performance, particularly in previously unseen, open-world settings. Human studies further reveal that the MLLM-as-a-judge variant produces more perceptually coherent placements than metric-driven approaches, highlighting a gap between standard evaluation metrics and human visual judgment.
Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is absent, large vision-language models usually address this task. We ask whether the same capability is available far more cheaply, from representations already learned by a...
K. Gupta, Ahmed Rafi Hasan, Md. Mahfuzur Rahman et al.· 0 citations
Under unseen distribution shifts, the OCWM with well-bound slots is more robust overall than the end-to-end trained scene-centric LeWM, while DINO-WM, built on similar frozen pretrained features, remains competitive -- pointing to pretrained features as a key contributor to robustness.
Shukrullo Nazirjonov, Sai Prasanna, Anna Manasyan et al.· 0 citations
OVIP-SG is presented, a unified framework for instance-preserving semantic mapping, functional scene partitioning, and language-guided small, fine-grained object retrieval that outperforms ConceptGraphs under a unified evaluation protocol on Replica.
Tianjing Hao, Hai-Yu Lan, Ang Li et al.· 0 citations
This work proposes RIG (Regional Image-prompt Generation), a novel training-free framework for multi-reference personalized generation that significantly outperforms state-of-the-art adapter methods in terms of both personalization fidelity and text-layout alignment.
Junhao Feng· Poster Volume 0008 The 2026...· 0 citations
Traditional Visual Relationship Detection (VRD) systems are strictly bound to predefined, closed-set vocabularies, limiting their practical application in real-world environments where object interactions are highly diverse. Moving towards an open-vocabulary setting (OV-VRD) provides more flexibility but introduces the...
N. Nguyen, Trung Tran· International Conference on...· 0 citations
An open-vocabulary semantic mapping pipeline is presented that integrates TALOS (TAgging–LOcation–Segmentation–Segmentation) with the probabilistic, instance-aware Voxeland framework and shows a better balance between map completeness, geometric clarity, semantic coherence, and instance separation.
Macoris Decena-Gimenez, Pepe Ojeda, J. Ruiz-Sarmiento et al.· Robotics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.