Skip to content
Preprint

Robostral Navigate

Abdelaziz Bounhar Abhijeet Somani Aditi Kabra Adrian Valente Adrien Petralia Adrien Sade Alan Jeffares Albert Q. Jiang А. В. Тимашов Alexandre Cahill Alexandre Gavaudan Alexandre Laval Alexandre Sablayrolles Amelie Heliou Amos You Andre Jonasson Andrew Bai Andrew Ehrenberg Andrew Zhao Angele Lenglemetz Anmol Agarwal Antonia Calvi Arata Suzuki Arjun Majumdar A. Fournier Artjom Joosen Avinash Sooriyarachchi A. Akkus Aysenur Karaduman Baptiste Bout Baptiste Roziere Baudouin De Monicault Benjamin Holzschuh Benjamin Lefaudeux Benjamin Tibi Bernhard Stadlbauer Blazej Osinski Camille Le Scao Chao Yu Charlotte Cronjager Chen Sun Chris Bamford Christian Wallenwein Christophe Renaudin Clemence Lanfranchi Corentin Barreau Corentin Sautier Cristiana-Diana Diaconu Cyprien Courtot Daniel Marczak Darius Dabert Diego de Las Casas Dominik Nuss Dylan Rubini Dzmitry Soupel Elizaveta Demyanenko Elliot Chane-Sane Emilien Fugier Emmanuel Gottlob E. Aas Étienne Goffinet Etienne Millon Eujeong Choi Fabian Paischer Fabian Schlager Faruk Ahmed Federico Baldassarre Filip Szatkowski Florian Wiesner Gabrielle Berrada Gaetan Ecrepont Gaétan Lepage Gaspard Blanchet Gaspard Donada-Vidal Gauthier Delerce Gauthier Guinet G. Hayes Georgii Novikov Giada Pistilli Gianluca Galletti Guillaume Breton Guillaume Kunsch Guillaume Lample Guillaume Martin Guillaume Raille Gunjan Dhanuka Gunshi Gupta Hang-Cheng Zhou Harshil Shah Hasan Furkan Vural Hedi Hadiji Hope McGovern Hugo Cisneros Hugo Thimonier Indraneel Mukherjee Ivan Cuevas Salazar Jacques Sun Jan Ludziejewski Jason Rute Jean Quentin Jean-Hadrien Chabran Jean-Malo Delignon Jie Zhang Joachim Studnia Joep Barmentlo J. Brandstetter John Harvill Jonas Amar Jonas Schweizer Joséphine Delas Josselin Somerville Julien Denize Julien Tauran Kartik Khandelwal Khyathi Raghavi Chandu Kilian Tep Kush Jain Larissa Laich Laura Calem Laurence Aitchison Laurent Callot Laurent Fainsin Leo Cotteleer Leonard Blier Lin Zhao L. Martin Louis Serrano Lucile Saulnier Ludovic Ho Fuh Luis Montero Maarten Buyl Manon Chossegros Marcin Mozejko Margaret Jennings Markus Hennerbichler M. Alexandre Mathieu Poir'ee Mathieu Schmitt Mathilde Guillaumin Matthieu Andre Matthieu Dinot Matthieu Futeral Maurits Bleeker Mauro Comi Max Mynter Maxim Berman Maxime Darrin Maxime Louis Maximilian Augustin Maximilian R. Müller Melina Jingting Laimon Mert Unsal Mia Chiquier Michael Pilcer Michal Pietruszka Michal Zajac Mikhail Biriuchinskii M. Pham Minsoo Kang Morgane Rivière Namit Katariya Nathan Grinsztajn Nathan Simpson Neeraj Aggarwal Neha Gupta Ola Mysiak Oliver Leicht Olivier Bousquet Olivier Duchenne Parag Jain Patricia Wang Patrick Blies Patrick von Platen Paul Jacob Paul Wambergue Paula Kurylowicz P. K. Reddy Pavel Kuksa Philippe Pinel Philomene Chagniot Pierre Stock Pierre-Andre Savalle Piotr Milos Prateek Gupta Pravesh Agrawal Quentin Desreumaux Quentin Torroba Quercus Hernandez Ram Ramrakhya Randall Isenhour Ranjit Parva Raul Perez Pelaez Reinhard Sonnleitner Remi Delacourt Richard Kurle Rishi Shah Rob Romijnders Rohin Arora Romain Sauvestre Roman Soletskyi R. Millner Rupert Menneer Sagar Vaze Samuel Barry Samuel Belkadi Samuel Humeau Sanchit Gandhi Sandeep Subramanian Sarthak Mittal Saskia Adaime S. Cha Sebastian Kaltenbach Shashwat Dalal Shashwat Verma S. Waly Shrimai Prabhumoye Siddhant Waghjale Siddharth Gandhi Simon Lepage Simon Sorg Soham Ghosh Sophie Marbach Srijan Mishra Stanislas Lange Steve Hong Sumukh Aithal Szymon Antoniak Tarun Kumar Vangani Teven Le Scao Théo Cachet Thibaut Lavril Thomas Chabal Thomas Coste Thomas Defard Thomas Foubert Thomas Robert Thomas Wang Tian-Yu Zhang Tim Lawson Timothee Lacroix Tobias Kronlachner Tom Bewley Tom Edwards Tomas Hodan Tuhin Das Tyler Wang Ulrick Ble Umar Jamil Umberto Tomasini Valentin Mace V. Phụng Vedant Nanda Victor Jouault Victor Letzelter Victor Paltz Victor Poucheret Vincent Maladiere Vincent Pfister Virgile Richard Vladislav Bataev Wassim Bouaziz Wen-Ding Li William Havard William Marshall Xinghui Li Xing-Ran Guo Xin-Yue-An-Nie Yang Yann Dreze Y. El Ouahidi Yassir Bendou Yi-Han Wang Yi-Mu Pan Yves Martin Des Taillades Zaccharie Ramzi Zheng Xu Zsofia Csakany
Jul 2026 · 2 citations · ⚡ 1 influential · 25 references
Computer Science

TL;DR

Robostral Navigate, an 8B vision-language model built around this scalability objective, is introduced, which consumes only a stream of monocular RGB images - the most ubiquitous sensor across robotic platforms and predicts waypoints by pointing to the next target location in the current camera view.

Abstract

Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployment cost. We introduce Robostral Navigate, an 8B vision-language model built around this scalability objective. The model consumes only a stream of monocular RGB images - the most ubiquitous sensor across robotic platforms and predicts waypoints by pointing to the next target location in the current camera view. Operating purely in image space, rather than robot-specific coordinates, makes the policy naturally robust to changes in camera intrinsics and scene scale, enabling deployment across wheeled, legged, and aerial robots without recalibration. We generate 2.4 million trajectories across 350k simulated scenes to reduce the reliance on real-world data collection and scale easily. We further introduce a prefix-caching training recipe that packs entire episodes into single training sequences, reducing training tokens by 22x and cutting training time from months to days. A tree-based attention mask prevents conditioning on previous ground-truth actions, encouraging visually grounded action prediction, and reinforcement learning is used to further improve exploration and recovery capabilities. On the Room-to-Room and Room-Across-Room in Continuous Environments (R2R-CE and RxR-CE) benchmarks, Robostral Navigate sets a new state of the art. On R2R-CE, it achieves a 77.4% success rate, surpassing the best monocular method by 10.5 points and the strongest depth- or multi-camera system by 5.3 points despite using only a single RGB camera. On RxR-CE, it reaches 75.1% success rate, outperforming all monocular baselines.

View source

Similar papers

Preprint Sep 2026

Multi-Task Visual Perception Network with LLM Conditioning for Autonomous Navigation

Long-term navigation for service robots faces crit- ical challenges like the accumulation of odometry drift and sensor error, which progressively degrade 2D maps and renders traditional path planning algorithms (e.g., A*, RRT*, DiPPer, ViT-A*) ineffective over time. To address this, we propose a user-friendly, interact...

Praveen Kumar, K. Guruprasad, Tushar Sandhan · 0 citations
Preprint Sep 2026

Robo-Harness K1: Harnessing Robot-Use Agents via Perception Augmentation

Robo-Harness K1 is introduced, a robot-use agent (RUA) framework that exposes perception as tools that makes 3D geometry accessible without changing the VLM architecture or training a depth encoder, and suggests that perception-augmented RUAs offer a promising route to sample-efficient, generalizable robotic policies t...

Ze-Xi Li, Ye-Hang Zhang, Wen-Qian Li et al. · 0 citations
Open access Aug 2026

Kitchen robotic manipulation utilizing foundation models

A modular perception pipeline for household manipulation tasks, with a focus on dishware handling in kitchen environments, that integrates open-vocabulary object detection, multi-view segmentation, instance-aware 3D reconstruction, and a 2D-3D feature fusion strategy for 6D pose estimation and grasp planning is present...

Myung-Hwan Jeon, Sankalp Yamsani, Joohyung Kim · 0 citations
Preprint Sep 2026

RoboFolDeX: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable Objects

Embodied AI, including vision-language-action and world-action models, must operate reliably in the physical world. Yet methods that perform well in simulation can degrade substantially on real robots, especially in long-horizon deformable-object manipulation, where policies must track changing states and execute relia...

Chen-Huan Liu, Yi Xu, Feng Wu et al. · 0 citations
Preprint Sep 2026

RoboMP-DINOv2: Prompts, Not Filters for Robust Robot Manipulation

Robot manipulation policies must generalize across visual shifts while preserving scene context relevant to action. General-purpose vision encoders are not tailored to visuomotor control, while object-centric approaches often use segmentation masks as hard filters that discard potentially useful context. We propose Rob...

Han Qi, Heng Yang · 0 citations
#machine learning Preprint Sep 2026

Generalizable Robotic Insertion with World Models

Robotic assembly in high-mixture settings requires adaptable systems that can handle diverse parts, yet current approaches typically rely on policies specialized to each insertion task. Although this can reach high success rates, it makes the process of deploying systems for new problems tedious and time consuming. We...

Nicklas Hansen, Iretiayo Akinola, Yi-Jie Guo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.