As the “power heart” of Urban Rail Transit (URT), the Traction Power Supply System (TPSS) is designed to provide sufficient traction power for trains. However, TPSS faults often lead to a Unidirectional Power Supply Shortage (UPSS), where the total power demand for trains is restricted within a local area. To mitigate the impact of UPSS on train operations, this paper proposes a distributed train rescheduling approach. The global problem is decomposed into three sub-problems through geographical and temporal partitioning, with tailored rescheduling models developed for each. For under-supplied area, a cooperative control model is developed to optimize train control strategies while maximizing regenerative power utilization. In well-supplied areas, a two-stage Train Timetable Rescheduling (TTR) model is proposed to maximize line capacity utilization and minimize the total waiting time of passengers using various rescheduling measures. Based on the initial results of the first two sub-problems, a feedback adjustment mechanism ensures global feasibility by synchronized times and the corresponding train trajectories in the cooperative control results. Finally, the TTR model is introduced for the recovery period to restore planned train operations. Given the safety-critical nature of train operations, a Hybrid-Shield Deep Q-Networks (HSDQN) algorithm is designed by enhancing the classical DQN with multiple protection mechanisms. Specifically, a preemptive shield enforces minimum headway constraints, while a post-posed shield addresses power supply capacity constraints. In addition, transition experiences with penalty terms are stored in the replay buffer to reduce the likelihood of unsafe actions. Finally, two numeral case studies based on Beijing Metro Yizhuang Line are presented to validate the proposed approach. Results show that HSDQN outperforms three reinforcement learning algorithms (CMB-IDQN, PDQN and SDQN) by up to 12.8%, 9.0%, and 2.6%, respectively. Compared to the previous TTR approach based on maximum traction power, the proposed distributed train rescheduling approach achieves performance improvements in both line capacity utilization and passenger waiting time.
Wang, X., D'Ariano, A., Su, S., Tang, T., Yu, K., Yan, M. (2026). Distributed train rescheduling considering regenerative power during unidirectional power supply shortages via safe reinforcement learning. TRANSPORTATION RESEARCH. PART C, EMERGING TECHNOLOGIES, 194 [10.1016/j.trc.2026.105984].
Distributed train rescheduling considering regenerative power during unidirectional power supply shortages via safe reinforcement learning
D'Ariano, Andrea;
2026-01-01
Abstract
As the “power heart” of Urban Rail Transit (URT), the Traction Power Supply System (TPSS) is designed to provide sufficient traction power for trains. However, TPSS faults often lead to a Unidirectional Power Supply Shortage (UPSS), where the total power demand for trains is restricted within a local area. To mitigate the impact of UPSS on train operations, this paper proposes a distributed train rescheduling approach. The global problem is decomposed into three sub-problems through geographical and temporal partitioning, with tailored rescheduling models developed for each. For under-supplied area, a cooperative control model is developed to optimize train control strategies while maximizing regenerative power utilization. In well-supplied areas, a two-stage Train Timetable Rescheduling (TTR) model is proposed to maximize line capacity utilization and minimize the total waiting time of passengers using various rescheduling measures. Based on the initial results of the first two sub-problems, a feedback adjustment mechanism ensures global feasibility by synchronized times and the corresponding train trajectories in the cooperative control results. Finally, the TTR model is introduced for the recovery period to restore planned train operations. Given the safety-critical nature of train operations, a Hybrid-Shield Deep Q-Networks (HSDQN) algorithm is designed by enhancing the classical DQN with multiple protection mechanisms. Specifically, a preemptive shield enforces minimum headway constraints, while a post-posed shield addresses power supply capacity constraints. In addition, transition experiences with penalty terms are stored in the replay buffer to reduce the likelihood of unsafe actions. Finally, two numeral case studies based on Beijing Metro Yizhuang Line are presented to validate the proposed approach. Results show that HSDQN outperforms three reinforcement learning algorithms (CMB-IDQN, PDQN and SDQN) by up to 12.8%, 9.0%, and 2.6%, respectively. Compared to the previous TTR approach based on maximum traction power, the proposed distributed train rescheduling approach achieves performance improvements in both line capacity utilization and passenger waiting time.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


