Gemini Robotics 2 brings whole body intelligence to robots
Google DeepMind Blog· deepmind.google· July 28, 2026
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
This survey reviews representative Transformer-based autonomous driving models and organizes them by task role, sensing configuration, and architectural design and analyzes how efficiency constraints reshaping model design choices in practice affects deployability, robustness, and safety.
Experiments on simulated and real-world benchmarks demonstrate that SCALE improves state-of-the-art VLAs and outperforms existing TTS methods while maintaining single-pass efficiency.
Hyeonbeom Choi, Daechul Ahn, Youhan Lee et al.· arXiv.org· 3 citations
A real-time semantic navigation framework for Unmanned Aerial Vehicles (UAVs) focused on improving time efficiency in the Object Goal Navigation (ObjectNav) task, using a Large Language Model that interprets user-provided natural language instructions and performs semantic reasoning over detected objects and spatial context to prioritize high-probability search regions.
Marin Maletic, Marijana Peti, T. Petrović et al.· European Conference on Mobil...· 3 citations
Mask2Real-WM is presented, a two-stage action-conditioned world model for dexterous manipulation that decouples pixel prediction into a dynamics model and a rendering model that shows that mask conditioning and simulation pretraining are both required for per-DoF action controllability across all 23 degrees of freedom.
Riccardo Feingold, Davide Liconti, Chenyu Yang et al.· 1 citation