Sep 2026· Communications in Transportation Research· 0 citations
Autonomous Vehicle Technology and Safety
TL;DR
This survey examines RL-based AD in modular and end-to-end pipelines and relates reported methods to task formulation and deployment evidence and examines deployment barriers, including safety, Sim2Real generalization, data efficiency, computation, embodied alignment, and evaluation readiness.
Abstract
Reinforcement learning (RL) is being studied for autonomous driving (AD), but its value depends on the role it plays in a task, the action interface, the evaluation protocol, and the evidence from deployment. This survey examines RL-based AD in modular and end-to-end pipelines and relates reported methods to task formulation and deployment evidence. It maps safe RL, offline RL, model-based RL, and PPO/GRPO-style fine-tuning to maneuver selection, continuous control, world modeling, and VLM/VLA-based driving. It also reviews simulators, datasets, RL platforms, and VLA benchmarks, with attention to reward design, observation space, traffic complexity, and open-loop versus closed-loop evaluation. The survey then examines deployment barriers, including safety, Sim2Real generalization, data efficiency, computation, embodied alignment, and evaluation readiness. RL and VLM/VLA-based methods have shown promise, but current evidence is insufficient to support reliable real-world deployment: many reported results come from restricted scenarios and depend on engineered rewards or simulator assumptions.
A comprehensive review of approximately 100 peer-reviewed studies that apply RL in the CARLA simulator shows that model-free RL overwhelmingly dominates the field, accounting for more than 80% of existing studies, with DQN, PPO, and SAC being the most frequently adopted algorithms.
A coherent map of the rapidly expanding landscape of visual RL is provided to provide researchers and practitioners with a coherent map of the rapidly expanding landscape of visual RL and to highlight promising directions for future inquiry.
In complex settings like smart manufacturing and human-robot teamwork, robots face internal disturbances and external uncertainties. Traditional control methods depend on accurate models and manual parameter tuning, leading to complex adjustment, weak dist urbance rejection, and poor generalization. Deep reinforcement...
. Quadrupedal robots exhibit strong mobility in complex environments where wheeled platforms often perform poorly, but their control remains difficult. In recent years, reinforcement learning (RL) has received growing attention in quadrupedal locomotion, as it supports direct policy optimization without relying entirel...
Chen Chang· Proceedings of the 3rd Inter...· 0 citations
CL4AD is presented, the first integration of curriculum learning into batched autonomous driving simulators by framing scenario selection as an unsupervised environment design problem, and utility functions that shape curricula based on success rates and the realism of the agent's behavior are introduced, in addition t...
Cevahir Koprulu, D. Paz, Feng Tao et al.· 1 citation
Reinforcement learning (RL) is increasingly used for real-time control of complex dynamical systems, but its practical performance must be evaluated under hardware constraints that are often simplified in simulation. This paper presents a comprehensive review of published RL-based control studies using the Quanser Aero...
Ghulam E. Mustafa Abro, S. Memon, Jawad Tanveer· Electronics· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Martin Trust Center Managing Director Bill Aulet introduces Dear Dreamer, a free platform for middle and high school students who want to learn about entrepreneurship.
Microsoft Research Blog· microsoft.comSep 30, 2026
Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduSep 25, 2026