SAGE improves robustness in novel and safety-critical scenarios, reduces safety violations, and maintains task performance comparable to strong baseline policies, suggesting that agents can improve after initial training by estimating the limits of their competence, requesting guidance when needed, and learning selectively from rare high-value events.
Abstract
Learning-based autonomous driving (AD) systems can perform reliably in familiar conditions, yet rare distribution shifts and long-tail events remain a major source of abrupt failure. A central limitation is that most agents learn primarily from passive experience and lack mechanisms to estimate when their competence is insufficient, seek timely assistance, and convert safety-critical encounters into targeted improvement. Here we present self-aware guided exploration (SAGE), an active learning framework for post-training adaptation in AD. SAGE learns a predictive world model that generates two online intrinsic signals: fear, which estimates short-horizon predictive risk and model uncertainty, and curiosity, which measures novelty through prediction error. Curiosity adaptively calibrates the intervention threshold for fear, allowing the agent to regulate risk in a context-dependent manner. When predicted fear exceeds this adaptive threshold, the agent transfers control to an expert or fallback policy and uses the resulting takeover trajectories for focused imitation learning. In parallel, fear is integrated into policy optimization and evaluation as a safety-oriented constraint to reduce performance regressions during adaptation. We evaluate SAGE in simulated route-transfer tasks, Waymo-based logged driving scenarios, CARLA occlusion hazards, and real-world mobile robot navigation tests. Across these settings, SAGE improves robustness in novel and safety-critical scenarios, reduces safety violations, and maintains task performance comparable to strong baseline policies. These results suggest that agents can improve after initial training by estimating the limits of their competence, requesting guidance when needed, and learning selectively from rare high-value events.
Autonomous driving has the potential to greatly enhance traffic efficiency, and its effectiveness depends on robust decision-making in complex real-world environments. As an emerging technique, Deep Reinforcement Learning (DRL) is expected to address this requirement. However, most existing general-purpose DRL methods...
Rui Guo, Xin-Yu Li, Zhong-Hao Fu et al.· IEEE Open Journal of Intelli...· 0 citations
Imitation learning (IL) teaches autonomous driving models what an expert does, but not why that action is safe or optimal. This critical gap arises because a single expert trajectory cannot illuminate the broader solution space: a complex performance landscape with multiple, distinct solutions and sharp “performance cl...
Skill adaptation frameworks based on reinforcement learning often require restrictive assumptions to maintain stability, such as fixed observations or tightly controlled exploration schedules. In cluttered and dynamic environments, however, unrestricted exploration can lead to unsafe behaviour and unstable learning, pa...
A. K. M. Nadimul Haque, Sheila Sutjipto, Marc G. Carmichael et al.· 0 citations
Offline safe reinforcement learning aims to learn constraint-satisfying policies from pre-collected datasets without online interaction, which is critical for safety-critical applications such as autonomous driving. However, offline datasets collected from heterogeneous sources often contain spurious correlations betwe...
Zi-Qian Wang, Zhen Zhang· IEEE/ASME International Conf...· 0 citations
End-to-end autonomous driving systems have demonstrated advantages over traditional modular systems. Despite this progress, these end-to-end systems still struggle to be deployed in real-world driving environments, as they inevitably encounter undertrained scenarios in which autonomous vehicles may take unsafe actions....
Daehyeok Kwon, Seung-Woo Seo, Sang-Hyun Lee· Proceedings of the Thirty-Fi...· 0 citations
This survey examines RL-based AD in modular and end-to-end pipelines and relates reported methods to task formulation and deployment evidence and examines deployment barriers, including safety, Sim2Real generalization, data efficiency, computation, embodied alignment, and evaluation readiness.
B. Shuai, Min Hua, Le-Tian Tao et al.· Communications in Transporta...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.