Skip to content
Review Open access

A Comprehensive Review of End-to-End Autonomous Driving: Architectures and Emerging Trends

Aug 2026 · Actuators · Vol 15, pp. 427 · 0 citations · 82 references

TL;DR

This review proposes a novel function-oriented taxonomy by categorizing architectures into perception-integrated and planning-integrated paradigms, and addresses critical challenges, particularly long-tail data scarcity and the deficiency in human-like decision-making.

Abstract

End-to-end autonomous driving is an emerging technology and a prominent research focus in both industry and academia. By integrating perception, localization, decision-making, and control into a single model, end-to-end systems aim to streamline the traditional modular pipeline while introducing new challenges in safety validation and interpretability. Unlike existing surveys that predominantly catalog algorithms, this review proposes a novel function-oriented taxonomy by categorizing architectures into perception-integrated and planning-integrated paradigms. Beyond the architectural dimension, the analysis delves into critical safety and interpretability, emphasizing the fundamental gap between theoretical design and the reliability required for real-world deployment. Industrial applicability is examined through real-world examples of data closed-loop workflows and simulation testing, addressing practical constraints in latency and computing resources. Finally, the review addresses critical challenges, particularly long-tail data scarcity and the deficiency in human-like decision-making and outlines future directions toward achieving robust autonomy.

Read PDF

Similar papers

Review Open access Jul 2026

Driving into the Future: A Comprehensive Survey of Autonomous Vehicle Technologies

Autonomous vehicles (AVs) represent a foundational cornerstone of future smart city transportation systems, offering the potential to eliminate human driving errors, reduce traffic fatalities by at least 40%, and optimize energy consumption. While AV technology is advancing rapidly toward highly automated driving (HAV) and driver-out Level 4 (L4) deployment, commercial realization remains hindered by significant technological uncertainties, soaring development costs, and strict safety-critical validation requirements. This provides a comprehensive, holistic survey of autonomous vehicle technology, bridging gaps in the existing literature by tracing its historical evolution to modern platforms equipped with LiDAR, radar, cameras, and  (V2X) communication. Beyond exploring essential architectural components, it evaluates the modern state of research by benchmarking the three dominant autonomous vehicle software architectures: modular pipelines, pure End-to-End (E2E), and hybrid systems across seven quantitative dimensions: planning quality, safety certification, latency, data efficiency, debuggability, Operational Design Domain (ODD) adaptation, and robustness. Comparative analysis of industry-standard benchmarks, including nuScenes and CARLA, reveals that while E2E and hybrid approaches achieve superior planning scores (88–91%) and lower collision rates (1.1–2.0%) than modular pipelines (84–86% planning; 3.95% collisions), pure E2E models lack a viable regulatory path for 2026 L4 deployment due to prohibitive validation mandates. Conversely, hybrid modular-E2E algorithms recover approximately 98% of E2E planning performance, drastically reduce debugging times to 2–6 hours, and retain compliance with ISO 26262:2018 ASIL-D safety standards. This work maps out the market landscape and identifies hybrid architectures as the current Pareto-optimal solution, balancing operational capability, safety assurance, and immediate regulatory feasibility for L3/L4 automation

Okpala Sanctus Emekumeh, Edje E. Abel, Ojugo A. Arnold · 0 citations
Review Jul 2026

The past, present and future of self-driving laboratories

The evolution of SDLs is charts their evolution from bespoke systems to interoperable platforms, highlighting challenges in scalability, generalizability and data provenance, and outlining pathways towards networked, trustworthy infrastructures enabling collective scientific superintelligence.

Richard B. Canty, M. Abolhasani · 0 citations
Review Open access 2026

Evolution of Scene Representation and Planning in Modular End-to-End Autonomous Driving: A Review

Modular end-to-end autonomous driving has emerged as a middle ground between classical modular stacks and opaque end-to-end pipelines, aiming to retain the trainability of end-to-end learning while improving interpretability through structured intermediate representations. This study presents a review of how modular end-to-end architectures evolve with respect to 1) scene representation and 2) planning design. We organize representative systems by their dominant scene representation — dense bird’s-eye-view (BEV) grids, vectorized/map-centric representations, and sparse token-based representations — and discuss how each choice shapes the design of prediction and planning heads, computational cost, and safety-critical failure modes. In parallel, we survey planning strategies used in these architectures, contrasting deterministic single-plan selection with probabilistic planning that explicitly models multi-modal futures and risk. Through a comparison of recent benchmarks and survey literature, we highlight recurring challenges, including error propagation across tasks, robustness under distribution shift, long-horizon reasoning, and evaluation under closed-loop interaction. Finally, we summarize open research directions for modular end-to-end driving, including scalable sparse representations, uncertainty-aware planning, representation–planning co-design, verification, and the integration of foundation and world models.

Arik Anjum Anik, Saeed S. Ba Hashwan, Mohd Haris Bin Md Khir et al. · 0 citations
Review Aug 2026

Planning-Oriented End-to-End Autonomous Driving: Architectures, Evaluation, and Emerging Paradigms

End-to-end autonomous driving has evolved from camera-to-control regression toward planning-oriented systems that use structured representations, trajectory-level outputs, and increasingly realistic evaluation protocols. This survey reviews this transition across behavior cloning, conditional imitation learning, privileged distillation, BEV and vectorized planning, unified perception-prediction-planning architectures, world-model-based planners, and vision-language-action systems. We argue that the key distinction in modern end-to-end driving is not whether intermediate representations are used, but whether they are learned, supervised, and evaluated to support safe, feasible, and route-compliant planning. To organize the literature, we synthesize existing methods along four axes: input representation, planning output, supervision signal, and evaluation protocol. We further examine the benchmark shift from open-loop trajectory matching to closed-loop simulation, non-reactive real-log evaluation, long-tail testing, and human-preference-aware metrics. Our analysis highlights that architectural progress is difficult to interpret without benchmark-consistent evaluation, and that displacement-based open-loop metrics alone provide limited evidence for safe and human-aligned driving. We conclude with open challenges in uncertainty-aware planning, learner-expert mismatch, runtime safety assurance, language-action grounding, world-model validation, and reproducible benchmarking.

Yanchen Guan, Xingcheng Liu, Bin Rao et al. · 0 citations
Book Open access Jul 2026

Agents in the Wild: Where Research Meets Deployment

Through applied case studies in pharmaceutical discovery and financial systems, common design patterns that make agentic systems successful are analyzed, and practical mitigation strategies for failure modes are discussed, such as verification pipelines, fallback mechanisms, and human-in-the-loop supervision.

Grace Hui Yang, P. Venkit, Hooman Sedghamiz et al. · 0 citations
Preprint Jul 2026

A Deployed Hybrid Vehicle-in-the-Loop Platform for Validating Cooperative Perception

European safety regulation now permits a large share of automated-driving homologation evidence to be produced virtually, provided a validated physical-virtual facility generates it. We present a deployed hybrid Vehicle-in-the-Loop (ViL) platform that couples a real instrumented vehicle with a CARLA-based digital twin (DT) through a V2X message pipeline, and we report its first integrated operation on a public-road-representative test track. A real vehicle streams ETSI-compliant CAM/CPM messages into the DT, where a GPU-accelerated Cooperative Perception (CP) module fuses them into a probabilistic occupancy grid during scenario runtime. We demonstrate the platform on a multi-vehicle double T-intersection scenario, characterise the CP workload across nominal, rain and night conditions and five localization-noise levels, and discuss the platform's current architectural limits and the engineering targets they define. The results show that CP substantially widens field-of-view (FoV) coverage and improves occupied-cell recall, and that beyond a moderate localization-noise threshold, positioning uncertainty, and not weather, becomes the dominant error source. We outline the platform's trajectory toward a Mediterranean operational design domain (ODD) testing service.

A. Bolovinou, Giorgos Hadjipavlis, M. Antonopoulos et al. · 0 citations