Skip to content
Preprint

Self-Aware Active Learning Enables Continual Improvement in Autonomous Driving

Aug 2026 · 0 citations · 42 references
Computer Science

TL;DR

SAGE improves robustness in novel and safety-critical scenarios, reduces safety violations, and maintains task performance comparable to strong baseline policies, suggesting that agents can improve after initial training by estimating the limits of their competence, requesting guidance when needed, and learning selectively from rare high-value events.

Abstract

Learning-based autonomous driving (AD) systems can perform reliably in familiar conditions, yet rare distribution shifts and long-tail events remain a major source of abrupt failure. A central limitation is that most agents learn primarily from passive experience and lack mechanisms to estimate when their competence is insufficient, seek timely assistance, and convert safety-critical encounters into targeted improvement. Here we present self-aware guided exploration (SAGE), an active learning framework for post-training adaptation in AD. SAGE learns a predictive world model that generates two online intrinsic signals: fear, which estimates short-horizon predictive risk and model uncertainty, and curiosity, which measures novelty through prediction error. Curiosity adaptively calibrates the intervention threshold for fear, allowing the agent to regulate risk in a context-dependent manner. When predicted fear exceeds this adaptive threshold, the agent transfers control to an expert or fallback policy and uses the resulting takeover trajectories for focused imitation learning. In parallel, fear is integrated into policy optimization and evaluation as a safety-oriented constraint to reduce performance regressions during adaptation. We evaluate SAGE in simulated route-transfer tasks, Waymo-based logged driving scenarios, CARLA occlusion hazards, and real-world mobile robot navigation tests. Across these settings, SAGE improves robustness in novel and safety-critical scenarios, reduces safety violations, and maintains task performance comparable to strong baseline policies. These results suggest that agents can improve after initial training by estimating the limits of their competence, requesting guidance when needed, and learning selectively from rare high-value events.

View source

Similar papers

Open access 2026

Adaptive Action-Constraint Safe Driving Decision Control Algorithm Based on Deep Reinforcement Learning

Autonomous driving has the potential to greatly enhance traffic efficiency, and its effectiveness depends on robust decision-making in complex real-world environments. As an emerging technique, Deep Reinforcement Learning (DRL) is expected to address this requirement. However, most existing general-purpose DRL methods...

Rui Guo, Xin-Yu Li, Zhong-Hao Fu et al. · 0 citations
Nov 2026

Consequence Learning for Trajectory Planning in Autonomous Driving

Imitation learning (IL) teaches autonomous driving models what an expert does, but not why that action is safe or optimal. This critical gap arises because a single expert trajectory cannot illuminate the broader solution space: a complex performance landscape with multiple, distinct solutions and sharp “performance cl...

Yi-Xuan Fan, Yali Li, Shengjin Wang · 0 citations
Preprint Sep 2026

Safety-aware Skill Adaptation for Reinforcement Learning in Dynamic Environments

Skill adaptation frameworks based on reinforcement learning often require restrictive assumptions to maintain stability, such as fixed observations or tightly controlled exploration schedules. In cluttered and dynamic environments, however, unrestricted exploration can lead to unsafe behaviour and unstable learning, pa...

A. K. M. Nadimul Haque, Sheila Sutjipto, Marc G. Carmichael et al. · 0 citations
Conference Aug 2026

Safe Offline Reinforcement Learning for Autonomous Driving via Causal Risk Features

Offline safe reinforcement learning aims to learn constraint-satisfying policies from pre-collected datasets without online interaction, which is critical for safety-critical applications such as autonomous driving. However, offline datasets collected from heterogeneous sources often contain spurious correlations betwe...

Zi-Qian Wang, Zhen Zhang · 0 citations
Conference Open access Sep 2026

Self-Improving Autonomous Vehicles via Real-World Reinforcement Learning

End-to-end autonomous driving systems have demonstrated advantages over traditional modular systems. Despite this progress, these end-to-end systems still struggle to be deployed in real-world driving environments, as they inevitably encounter undertrained scenarios in which autonomous vehicles may take unsafe actions....

Daehyeok Kwon, Seung-Woo Seo, Sang-Hyun Lee · 0 citations
#reinforcement learning Review Open access Sep 2026

Recent Advances of Reinforcement Learning Algorithms for Autonomous Driving System

This survey examines RL-based AD in modular and end-to-end pipelines and relates reported methods to task formulation and deployment evidence and examines deployment barriers, including safety, Sim2Real generalization, data efficiency, computation, embodied alignment, and evaluation readiness.

B. Shuai, Min Hua, Le-Tian Tao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.