Evolution of Scene Representation and Planning in Modular End-to-End Autonomous Driving: A Review
Modular end-to-end autonomous driving has emerged as a middle ground between classical modular stacks and opaque end-to-end pipelines, aiming to retain the trainability of end-to-end learning while improving interpretability through structured intermediate representations. This study presents a review of how modular end-to-end architectures evolve with respect to 1) scene representation and 2) planning design. We organize representative systems by their dominant scene representation — dense bird’s-eye-view (BEV) grids, vectorized/map-centric representations, and sparse token-based representations — and discuss how each choice shapes the design of prediction and planning heads, computational cost, and safety-critical failure modes. In parallel, we survey planning strategies used in these architectures, contrasting deterministic single-plan selection with probabilistic planning that explicitly models multi-modal futures and risk. Through a comparison of recent benchmarks and survey literature, we highlight recurring challenges, including error propagation across tasks, robustness under distribution shift, long-horizon reasoning, and evaluation under closed-loop interaction. Finally, we summarize open research directions for modular end-to-end driving, including scalable sparse representations, uncertainty-aware planning, representation–planning co-design, verification, and the integration of foundation and world models.