Skip to content
#robotics Preprint

Keep the Future, Drop the Rollout: RIFT for World Action Models

Aug 2026 · 1 citation · 50 references
Computer Science

Abstract

World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on 40 simulated robotic manipulation tasks, paired closed-loop interventions show that blocking access to the future cache or reassigning its values changes execution and reduces success. Yet in the evaluated co-denoising settings, reusing one fixed final-clean key/value (K/V) cache throughout action denoising nearly preserves unmodified execution, with $1.7$--$1.9$ cm end-effector average displacement error. Obtaining this cache still requires iterative video generation. We therefore propose RIFT (Rollout-free Imagination via Future Tokens), which uses learned anticipation tokens to construct a complete future K/V cache in one backbone pass. On LIBERO, RIFT achieves $98.8\%$ overall success, outperforming all evaluated rollout-based methods while yielding a $3.1$--$9.2\times$ inference speedup. Without further training, it achieves $81.1\%$ overall success on the out-of-distribution LIBERO-Plus benchmark, a $+9.7$ percentage-point improvement over the strongest evaluated baseline. On RoboTwin, it achieves $92.9\%$ and $92.6\%$ success on clean and randomized scenes, respectively, the highest among the evaluated methods. On real-world manipulation tasks, RIFT achieves $45.3\%$ average success, a $+6.0$ percentage-point improvement over Fast-WAM-Joint. These results support rollout-free future conditioning without iterative video generation at deployment.

View source

Similar papers

#artificial intelligence Review Apr 2023

Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A Survey

This survey reviews representative Transformer-based autonomous driving models and organizes them by task role, sensing configuration, and architectural design and analyzes how efficiency constraints reshaping model design choices in practice affects deployability, robustness, and safety.

J. Zhong, Zheng Liu, Xiangshan Chen · 21 citations
#artificial intelligence Preprint Sep 2026

Show-Harness: Just a VLM Agent Can Play Robots

Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to"play"robots through a compact semantic interface linking intent to action. Show-Harness exposes...

Yan-Zhe Chen, Ze-Chen Bai, Zhi-Jun Cao et al. · 16 citations · ⚡2

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation

Preliminary results indicate that a world-action model (WAM) post-trained from OmniDreams achieves strong performance on the Physical AI Autonomous Vehicles NuRec dataset, surpassing the VLA-based Alpamayo 1.5 research policy model while using only 1/5 the total parameters.

Aarti Basant, Amlan Kar, Despoina Paschalidou et al. · 14 citations · ⚡2
#artificial intelligence Open access May 2025

Building Intelligent Agents with Neuro-Symbolic Concepts

A concept-centric framework for building agents that can learn continually and reason flexibly across multiple domains and offers several advantages, including data efficiency, compositional generalization, continual learning, and zero-shot transfer.

Jia-Yuan Mao, Joshua B. Tenenbaum, Jia-Jun Wu · 14 citations · ⚡1
#artificial intelligence Preprint Feb 2026

VLANeXt Family: A Systematic Study of VLA Models from Core Recipes to Emerging Paradigms

This study systematically dissects design choices along three dimensions: foundational components, perception essentials, and action modeling perspectives and distill 12 key findings that together form a practical recipe for building strong VLA models, resulting in a simple yet effective model, VLANeXt.

Xiao-Ming Wu, Kang Liao, Yi-Hang Luo et al. · 12 citations · ⚡3

Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring

This paper proposes Hide-and-Seek, a framework that formulates VLA failure detection as a coarsely supervised learning problem that achieves state-of-the-art multi-task failure detection performance with a practical accuracy--timeliness trade-off under conformal prediction, and generalizes well to both seen and unseen...

S. Park, Wendi Li, Changdae Oh et al. · 8 citations

Related blog posts

Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.