Skip to content
#reinforcement learning Review Open access

Recent Advances of Reinforcement Learning Algorithms for Autonomous Driving System

Sep 2026 · Communications in Transportation Research · 0 citations
Autonomous Vehicle Technology and Safety

TL;DR

This survey examines RL-based AD in modular and end-to-end pipelines and relates reported methods to task formulation and deployment evidence and examines deployment barriers, including safety, Sim2Real generalization, data efficiency, computation, embodied alignment, and evaluation readiness.

Abstract

Reinforcement learning (RL) is being studied for autonomous driving (AD), but its value depends on the role it plays in a task, the action interface, the evaluation protocol, and the evidence from deployment. This survey examines RL-based AD in modular and end-to-end pipelines and relates reported methods to task formulation and deployment evidence. It maps safe RL, offline RL, model-based RL, and PPO/GRPO-style fine-tuning to maneuver selection, continuous control, world modeling, and VLM/VLA-based driving. It also reviews simulators, datasets, RL platforms, and VLA benchmarks, with attention to reward design, observation space, traffic complexity, and open-loop versus closed-loop evaluation. The survey then examines deployment barriers, including safety, Sim2Real generalization, data efficiency, computation, embodied alignment, and evaluation readiness. RL and VLM/VLA-based methods have shown promise, but current evidence is insufficient to support reliable real-world deployment: many reported results come from restricted scenarios and depend on engineered rewards or simulator assumptions.

Read PDF

Similar papers

Review Open access Sep 2025

A Comprehensive Review of Reinforcement Learning for Autonomous Driving in the CARLA Simulator

A comprehensive review of approximately 100 peer-reviewed studies that apply RL in the CARLA simulator shows that model-free RL overwhelmingly dominates the field, accounting for more than 80% of existing studies, with DQN, PPO, and SAC being the most frequently adopted algorithms.

Elahe Delavari, Feeza Khan Khanzada, Jaerock Kwon · 12 citations
Review Open access Aug 2026

Reinforcement Learning for Multimodal Foundation Models: A Survey

A coherent map of the rapidly expanding landscape of visual RL is provided to provide researchers and practitioners with a coherent map of the rapidly expanding landscape of visual RL and to highlight promising directions for future inquiry.

Weijia Wu, Chen Gao, Joya Chen et al. · 0 citations
#reinforcement learning Conference Open access Sep 2026

A Review of Robot Adaptive Control Driven by Deep Reinforcement Learning

In complex settings like smart manufacturing and human-robot teamwork, robots face internal disturbances and external uncertainties. Traditional control methods depend on accurate models and manual parameter tuning, leading to complex adjustment, weak dist urbance rejection, and poor generalization. Deep reinforcement...

Fu-Te Xiao · 0 citations
Conference Open access 2026

Reinforcement Learning for Quadrupedal Robot Control: Taxonomy, Sim-to-Real, Robustness, and Emerging Trends

. Quadrupedal robots exhibit strong mobility in complex environments where wheeled platforms often perform poorly, but their control remains difficult. In recent years, reinforcement learning (RL) has received growing attention in quadrupedal locomotion, as it supports direct policy optimization without relying entirel...

Chen Chang · 0 citations
Preprint Aug 2026

Scaling Curriculum Learning For Autonomous Driving

CL4AD is presented, the first integration of curriculum learning into batched autonomous driving simulators by framing scenario selection as an unsupervised environment design problem, and utility functions that shape curricula based on success rates and the realism of the agent's behavior are introduced, in addition t...

Cevahir Koprulu, D. Paz, Feng Tao et al. · 1 citation
Review Open access Sep 2026

Reinforcement Learning for Real-Time Control Using Quanser Platforms: A Structured Narrative Review

Reinforcement learning (RL) is increasingly used for real-time control of complex dynamical systems, but its practical performance must be evaluated under hardware constraints that are often simplified in simulation. This paper presents a comprehensive review of published RL-based control studies using the Quanser Aero...

Ghulam E. Mustafa Abro, S. Memon, Jawad Tanveer · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.