A novel Dual Q–Adaptive Weighting model that balances reward and cost through a primal–dual learning mechanism is proposed and achieves the highest route accuracy of 78.66%, outperforming Double Q-Learning, Q-Learning, and traditional A* and Dijkstra algorithms.
Abstract
Electric Vehicles (EVs) are emerging as sustainable alternatives to internal combustion engine vehicles; however, efficient route planning remains a major challenge due to limited driving range, sparse charging infrastructure, and variable energy consumption patterns. Traditional shortest-path algorithms, such as Dijkstra’s and A*, often fail to account for EV-specific factors, including charging station availability, connector compatibility, and energy constraints. This study presents a comprehensive EV route optimization framework that integrates reinforcement learning (RL) with graph-based methods. A novel Dual Q–Adaptive Weighting model that balances reward and cost through a primal–dual learning mechanism is proposed. The framework learns energy-aware routing strategies from historical navigation experience. The model is compared against standard RL approaches—Q-Learning and Double Q-Learning—as well as enhanced variants of A* and Dijkstra’s algorithms that incorporate charging density and time-penalty considerations. Real-world EV charging infrastructure data from the Alternative Fuels Data Center (AFDC) and Placekey datasets are used to construct a clustered navigation graph via DBSCAN. Experimental results across multiple intercity routes show that the proposed Dual Q–Adaptive model achieves the highest route accuracy of 78.66%, outperforming Double Q-Learning (76.27%), Q-Learning (77.52%), and traditional A* (74.26%) and Dijkstra (60.92%) algorithms. A* and Dijkstra with modifications, use fewer charging stops than traditional algorithms. The Improvised algorithms provide substantial improvements over their baseline counterparts. The results demonstrate that reinforcement learning integrated with graph-theoretic optimization can enable scalable, infrastructure-aware, and efficient EV route planning.
A hybrid optimization framework that combines greedy initialization with reinforcement learning to efficiently explore the charging station deployment problem is proposed and demonstrates stable performance across three evaluated deployment scenarios, indicating its potential applicability to increasingly complex charging infrastructure planning problems.
Electric Vehicles are thought to be among the best options for lowering gas emissions and oil consumption. EV users can from a charging station with a well-thought-out schedule and price plan. Advanced energy management techniques are required to guarantee the sustainable, dependable, and effective operation of charging infrastructure due to the quick rise in the usage regarding electric vehicle (EV). In light of recent research, current design restrictions as well as the erratic conduct of EV consumers make classic scheduling techniques, such as set costs for Time-of-Use (ToU), insufficient. Uncoordinated charging results Voltage instability caused by transformer overloading, also wasteful utilization of electricity from renewable as EV use rises. AI is becoming more widely acknowledged in a crucial facilitator of scalable, durable, effective EV charging facilities with intelligence. This paper proposes a hybrid AI-based architecture that integrates real-time Traffic pattern, distance of EV, arrival and departure time of EV state of charge as input. Through real-time monitoring and charge optimization, the EVCS enable intelligent EV charging. The AI framework employs a non-uniform Poisson process in order to dynamically assess user demand also enhances schedule of charging. While the charging demand of electric vehicles (EVs) is intrinsically heterogeneous, decentralized, and stochastic, the intermittent and weather-dependent nature of PV power results in considerable output uncertainty. The purpose of the EV Charging Grid Optimization is to facilitate research on AI-driven energy management for EV charging infrastructure. The two proposed optimization algorithm improves the operational effectiveness of EVCS. Using threshold value Reinforcement Learning make the real-time decision that dynamically schedule the EV. Threshold value is determined from customer preferences. The proposed technology demonstrates scalability, durability, and cost-effectiveness and provides a feasible substitute for upcoming metropolitan EV charging system.
Jose Devaraj, Daphni Paulphin J.· International journal of com...· 0 citations
In the era of the Internet of Things (IoT), coordinating connected electric vehicle (EV) charging scheduling to balance EV charging satisfaction, station profitability, and smart grid stability presents a complex multi-objective challenge. Existing Multi-Agent Reinforcement Learning (MARL) approaches often struggle with high-dimensional state spaces generated by massive IoT sensing data and conflicting stakeholder interests. This paper proposes a novel LLM-enhanced MARL framework that, for the first time, simultaneously optimizes the Grid, EVs, and Stations within a unified loop. By integrating Large Language Model (LLM), we address two critical bottlenecks: interpretable feature selection and adaptive multi-objective balancing. The LLM analyzes real-time IoT-collected environmental states to extract physically significant features and dynamically assigns weights to conflicting objectives-including profit, user satisfaction, and grid load-using semantic reasoning instead of complex manual tuning. Extensive experiments demonstrate that our framework significantly outperforms state-of-the-art baselines, achieving superior market efficiency while reducing training time by over 70%. This approach offers a scalable, transparent solution for efficient and sustainable IoT-enabled urban charging infrastructure management.
Yang Zhang, Lin-Dong Xie, Chong-Yu Wang et al.· 0 citations
Deploying heavy-duty electric trucks under real-world uncertainty is operationally challenging, particularly when multiple vehicles compete for limited public charging resources and face uncertain wait times. This research studies the Fixed-Route Vehicle Charging Problem and formulates it as a multistage stochastic program under charging congestion uncertainty. The delivery system is modeled as a discrete-event process triggered by physical route milestones, and charging congestion uncertainty is represented by a Markovian transition. To address the resulting mixed-integer structure, we apply stochastic dual dynamic programming (SDDP) as a tactical planning approach, while accommodating discrete vehicle dynamics within the convexity requirements. Computational experiments on a California logistics network showed that the proposed approach performed effectively across multiple geographically diverse delivery routes. Compared to a multistage stochastic integer programming benchmark, SDDP achieved a competitive solution quality while reducing training time as the number of route instances increased. Out-of-sample simulations further demonstrated that policies derived from SDDP remained robust under distributional shifts toward more congested scenarios. Overall, this study establishes SDDP as a tractable and scalable framework for generating high-quality operational policies for heavy-duty electric fleets under congestion uncertainty.
Ziyan Li, Nikolay Aristov, E. Dugundji· Transportation Research Reco...· 0 citations
These findings demonstrate that reinforcement learning is a promising and scalable alternative to conventional heuristic and metaheuristic approaches for capacitated routing problems, particularly in dynamic logistics environments that require rapid and adaptive decision making.
Audrey Ariij Sya'imaa HS, Muhamad Rizky Aulia, Siti Hadiaty Yuningsih· International Journal of Mat...· 0 citations
Vehicle-to-grid (V2G) technology lets a parked electric vehicle push power back to the network, so a fleet of cars can act as distributed storage. Many studies report large benefits from this, such as peak shaving and better use of renewable energy, but most describe their simulation only in words, name no test network, and compare a single charging behavior against a do-nothing case. It is then hard to separate what V2G delivers from what the control strategy delivers. This paper builds and compares dispatch strategies on one fully specified system: a price-following rule with no knowledge of the network, a stronger randomized time-of-use baseline, and a receding-horizon model predictive control (MPC) strategy that re-plans every hour with the distribution network’s per-line thermal limits and voltage bounds embedded directly in the optimization, not merely checked afterward. All run on the IEEE 33-bus feeder at low (10%), medium (30%), and high (50%) EV shares, with a full AC power flow solved every hour. The central result is a critical-penetration effect. At a 50% share, the naive price rule makes the system peak 23.7% worse than having no V2G at all, because the whole fleet reacts to one price signal at once, while the MPC cuts the peak by 18 to 28% in every case. Embedding the network limits removes the worst-line-loading side-effect a system-peak-only objective creates, bringing high-share worst-line loading down from 86.3% to 81.8% while preserving the full peak reduction. A workplace daytime-charging scenario shows renewable use becoming a real, separating metric (up to 2.8 MWh/day of EV demand met directly by rooftop solar, against zero for overnight charging), a quadratic wear cost smooths the profit-cycling frontier that a linear cost makes step-shaped, the controller is robust to forecast error up to 20%, and its solve time is set by network size rather than fleet size. Results are given as they came out of the model.
Muhammad Abdullah Bin Arif, S. Iqbal, Sanchari Deb· Energies· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.