Skip to content

Latency-Aware Residual DRL for Energy-Efficient Multi-RIS Assisted High-Speed Railway Communications

2026 · IEEE Transactions on Green Communications and Networking · Vol 10, pp. 4174-4189 · 0 citations · 35 references

Abstract

Maximizing energy efficiency in multiple Reconfigurable Intelligent Surface (RIS)-assisted High-Speed Railway (HSR) networks presents a formidable challenge due to the coupled dynamics of high-mobility Doppler effects, strict latency constraints, and the necessity for joint transmit power and beamforming control. Existing optimization approaches, such as genetic algorithms (GA) and Continuous relaxation methods, often suffer from limited practical effectiveness due to high computational latency, leading to beam aging–a critical misalignment between the optimized beam and the fast-moving train. Furthermore, these methods typically rely on fixed power allocation or saturated optimization strategies that fail to adapt to rapid channel fluctuations. To address these issues, we propose a novel Latency-aware Residual Twin Delayed Deep Deterministic Policy Gradient (TD3) Deep Reinforcement Learning (DRL) framework. Unlike standard opaque DRL, our agent integrates: 1) a clamped residual learning architecture that jointly fine-tunes geometric beamforming and transmit power; and 2) a zero-initialization strategy to eliminate unstable warm-up phases. We evaluate this framework under a 3GPP TR 38.901-compliant high-mobility channel model with spatially consistent, per-path Doppler generation and delay-affected channel state information, benchmarking against classical (Dinkelbach-SLSQP) and metaheuristic (GA) solvers as well as a predictive extended Kalman filter-based tracking baseline that isolates the benefit of latency prediction from that of learning-based control. Comprehensive Monte Carlo simulations, including a ±15% time-varying velocity profile that stress-tests robustness to both random fading and kinematic uncertainty, show that the proposed DRL agent consistently achieves the lowest optimality gap and the tightest confidence intervals of all methods, particularly at ultra-high speeds (500 km/h) with large-scale RIS arrays, through a power allocation strategy that intelligently backs off power under favorable channel conditions to maximize efficiency.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.