Empirical comparisons show that PUI-MAPPO (multi-agent proximal policy optimization) achieves the best performance among all PUI-enhanced variants, and ablation studies further validate the individual effectiveness of the PUI urgency mechanism, the dynamic threshold framework, and the adaptive reward function.
Abstract
This paper proposes a Priority–Urgency Index (PUI) mechanism to address the multi-objective conflict problem in electric vehicle charging scheduling at public fast-charging stations. The mechanism dynamically quantifies the charging urgency of each vehicle based on remaining dwell time, current state of charge (SoC), and target SoC, enabling differentiated service prioritization under resource scarcity. The PUI is systematically integrated into four multi-agent reinforcement learning (MARL) algorithms. To handle the time-varying number of vehicle entities caused by random arrivals and departures, the chargers are modeled as fixed agents; a partially observable Markov decision process (POMDP) is formulated, and a centralized training with decentralized execution (CTDE) architecture is adopted. On this basis, a state-aware dynamic threshold mechanism is introduced to distinguish urgency levels of charging tasks, and an adaptive reward function is designed to accommodate complex operating conditions. Empirical comparisons show that PUI-MAPPO (multi-agent proximal policy optimization) achieves the best performance among all PUI-enhanced variants. Under extreme supply–demand conditions—such as resource-scarce and heavy-traffic scenarios—PUI-MAPPO improves the target-SoC fulfillment rate and net revenue by up to 42.7% and 23.2%, respectively, and reduces the cumulative grid-limit exceedance by 22.8% to 47.3%, relative to the first-come, first-served (FCFS) baseline. Ablation studies further validate the individual effectiveness of the PUI urgency mechanism, the dynamic threshold framework, and the adaptive reward function.
A collaborative optimization framework based on multi-agent reinforcement learning is proposed for orderly charging at electric vehicle charging stations and coordinated interaction with the power grid, providing a technical reference for intelligent charging coordination under grid interaction and electromagnetic comp...
This paper proposes a novel LLM-enhanced MARL framework that, for the first time, simultaneously optimizes the Grid, EVs, and Stations within a unified loop by integrating Large Language Model (LLM).
Yang Zhang, Lin-Dong Xie, Chong-Yu Wang et al.· 0 citations
Electric vehicle (EV) cluster access is becoming a key control problem for urban intelligent transportation charging facilities, because concentrated arrivals at residential and public chargers can reshape feeder loading, driver waiting time, and voltage security at the same time. This paper proposes an intelligent rol...
Hao Wu, Qiu-Shi Xu, Xiao-Yun Wang et al.· International Conference on...· 0 citations
: Random vehicle arrivals, heterogeneous charging demands, and the station-level limit on aggregate electric vehicle (EV) charging power pose simultaneous challenges to event-driven ordered charging in terms of causal information constraints, charging economics, and aggregate-load coordination. Existing full-informatio...
Li-Xiao Wang, Jia-Qi Li, Hai-Feng Li et al.· Energy Engineering· 0 citations
A hybrid AI-based architecture that integrates real-time Traffic pattern, distance of EV, arrival and departure time of EV state of charge as input, and two proposed optimization algorithm improves the operational effectiveness of EVCS.
Jose Devaraj, Daphni Paulphin J.· International journal of com...· 0 citations
A hybrid optimization framework that combines greedy initialization with reinforcement learning to efficiently explore the charging station deployment problem is proposed and demonstrates stable performance across three evaluated deployment scenarios, indicating its potential applicability to increasingly complex charg...
A. Bousia· Sustainability· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026