Multi-modal planning is promising for autonomous driving by representing multiple plausible behaviors in ambiguous and long-tail scenarios. Existing methods mainly focus on improving trajectory multi-modality, enhancing trajectory representations, or reshaping the candidate distribution. Nevertheless, we identify a pro...
Automated parking commonly assumes marked slots and short approach maneuvers. Delivery and service vehicles, however, may need to reach an operator-specified pose in an irregular bounded environment from a distant start. Existing learning-based parking planners often rely on local observations, which can restrict long-...
Zihan Wang, Baixiang Huang, Yang Guan et al.· 0 citations
On-policy distillation (OPD) has emerged as an effective approach for large language model post-training, yet existing objectives face a trade-off between objective fidelity and optimization stability. Token-level OPD provides stable but local supervision, whereas sequence-level OPD captures future credit at the cost o...
Shi-Qi Liu, Ze-Yu He, Le-Tian Tao et al.· 2 citations
This survey examines RL-based AD in modular and end-to-end pipelines and relates reported methods to task formulation and deployment evidence and examines deployment barriers, including safety, Sim2Real generalization, data efficiency, computation, embodied alignment, and evaluation readiness.
B. Shuai, Min Hua, Le-Tian Tao et al.· Communications in Transporta...· 0 citations
This paper proposes exchange policy optimization (EPO), an algorithmic framework that achieves optimal policy performance with provably bounded safety guarantees and derives an upper bound on the required number of iterations and quantifies the gap between the obtained policy and the true optimum.
Jiaming Zhang, Yujie Yang, Hao-Ning Wang et al.· arXiv.org· 1 citation
An AIM framework based on residual-penalty variable splitting, which interprets momentum as a multiplier-like correction driven by the splitting residual, and RADAR, which combines relativistic adaptive geometry, decoupled residual correction, and second-order momentum filtering to improve the update direction and mome...
Zhi-Xin Ren, Yao Lyu, Cong-Rong Li et al.· 2 citations
A joint identifiability condition for controlled world models with Gaussian latent states with Gaussian latent states is presented, which consists of two coupled components: representation identifiability and transition identifiability, and it is proved that when this condition holds, minimizing the LeJEPA-style predic...
Xiangteng Zhang, Yang Guan, Bo Zhang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.