Skip to content

7 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Evaluation Is All You Need for Multi-Modal Autonomous Driving

Multi-modal planning is promising for autonomous driving by representing multiple plausible behaviors in ambiguous and long-tail scenarios. Existing methods mainly focus on improving trajectory multi-modality, enhancing trajectory representations, or reshaping the candidate distribution. Nevertheless, we identify a pro...

Ze-Yu He, Shi-Qi Liu, Ke Chen et al. · 0 citations
Preprint Aug 2026

NeuralParker: A Reinforcement Learning Planner for Irregular Parking Environments

Automated parking commonly assumes marked slots and short approach maneuvers. Delivery and service vehicles, however, may need to reach an operator-specified pose in an irregular bounded environment from a distant start. Existing learning-based parking planners often rely on local observations, which can restrict long-...

Zihan Wang, Baixiang Huang, Yang Guan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation

On-policy distillation (OPD) has emerged as an effective approach for large language model post-training, yet existing objectives face a trade-off between objective fidelity and optimization stability. Token-level OPD provides stable but local supervision, whereas sequence-level OPD captures future credit at the cost o...

Shi-Qi Liu, Ze-Yu He, Le-Tian Tao et al. · 2 citations
#reinforcement learning Review Open access Sep 2026

Recent Advances of Reinforcement Learning Algorithms for Autonomous Driving System

This survey examines RL-based AD in modular and end-to-end pipelines and relates reported methods to task formulation and deployment evidence and examines deployment barriers, including safety, Sim2Real generalization, data efficiency, computation, embodied alignment, and evaluation readiness.

B. Shuai, Min Hua, Le-Tian Tao et al. · 0 citations

Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning

This paper proposes exchange policy optimization (EPO), an algorithmic framework that achieves optimal policy performance with provably bounded safety guarantees and derives an upper bound on the required number of iterations and quantifies the gap between the obtained policy and the true optimum.

Jiaming Zhang, Yujie Yang, Hao-Ning Wang et al. · 1 citation
#machine learning Preprint Aug 2026

Momentum as Residual-Driven Multiplier Correction for Deep Learning Optimization

An AIM framework based on residual-penalty variable splitting, which interprets momentum as a multiplier-like correction driven by the splitting residual, and RADAR, which combines relativistic adaptive geometry, decoupled residual correction, and second-order momentum filtering to improve the update direction and mome...

Zhi-Xin Ren, Yao Lyu, Cong-Rong Li et al. · 2 citations
Jul 2026

On the Identifiability of Controlled World Models

A joint identifiability condition for controlled world models with Gaussian latent states with Gaussian latent states is presented, which consists of two coupled components: representation identifiability and transition identifiability, and it is proved that when this condition holds, minimizing the LeJEPA-style predic...

Xiangteng Zhang, Yang Guan, Bo Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.