Skip to main content
Article

Deep reinforcement learning–driven electric vehicle charging coordination for renewable-integrated intelligent energy districts

Abdellatif M. SadeqFaculty of Agricultural Mechanization, TIIAME National Research University, Kori Niyoziy 39, 100000, Tashkent, UzbekistanSaman AminianDepartment of Civil Engineering, College of Engineering, Cihan University-Erbil, Erbil, IraqRashed Abu HammourFaculty of Technical Education , Hourani Center for Applied Scientific Research (HCASR), Al-Ahliyya Amman University, Amman, JordanArti BadhoutiyaDepartment of Electrical Engineering, GLA University, Mathura, IndiaNouf Abd ElmunimDepartment of Electrical Engineering, College of Engineering, Princess Nourah bint Abdulrahman University, P.O. Box 84428, 11671, Riyadh, Saudi ArabiaMohamed ShabanPhysics Department, Faculty of Science, Islamic University of Madinah, P. O. Box: 170, 42351, Madinah, Saudi ArabiaSujeet KumarDivision of Research and Development, Lovely Professional University, Phagwara, Punjab, IndiaHusam RajabCollege of Engineering, Department of Mechanical Engineering, Najran University, King Abdulaziz Road, P.O Box 1988, Najran, Saudi ArabiaNarinderjit Singh Sawaran SinghFaculty of Data Science and Information Technology, INTI International University, Persiaran Perdana BBN, Putra Nilai, 71800, Nilai, Malaysia
2026en
ABI

Abstract

Coordinated electric-vehicle charging in renewable-integrated intelligent energy districts must respond within quarter-hourly operating windows while balancing uncertain photovoltaic generation, tariffs, mobility demand, and distribution-feeder limits. This study develops and evaluates a closed-loop hybrid framework in which a four-head Generalized Value Function Network constructs predictive states, asynchronous IMPALA coordinates ten district controllers, and an SOCP–DistFlow projection screens the aggregate dispatch against relaxed-model voltage and branch-loading constraints. The framework was tested on the IEEE 69-bus feeder using two years of operational records for ten districts and 70 electric vehicles. Within the reported episode histories, IMPALA reached a sustained 90%-of-peak reward threshold by episode 50, retained 99.6% of its best 100-episode reward, and exhibited a 3.0% post-convergence coefficient of variation. These statistics support method selection within the observed histories and do not establish between-seed superiority. The integrated pipeline completed a quarter-hourly schedule in approximately 4.8 min, compared with 26.9 min for the pure-optimization benchmark, representing an 82% reduction in runtime with an approximately 3% cost deviation. Sensitivity and joint-uncertainty analyses indicated stable mean performance over the tested envelopes, while severe joint conditions increased voltage, branch-loading, and upper-tail cost exposure. The results show that predictive-state learning, asynchronous coordination, and feeder-level conic screening can produce deadline-compatible schedules with a transparent cost–latency trade-off for the evaluated simulation. The evidence remains limited to the relaxed SOCP feeder model and reported histories; nonlinear AC replay, repeated seeded trials, hardware-in-the-loop testing, and controlled utility pilots remain necessary before field deployment.

Identifiers

Citations and references

Cited by 00 references