Aug 2026· Neural Networks· Vol 205 Pt C, pp.
109508
· 0 citations· 66 references
Medicine
TL;DR
This study proposes a deployment-oriented edge-cloud collaboration (ECC) framework integrated with a transient-aware predictive architecture, named FS-Attention, designed to balance transient responsiveness, engineering deployability, and decision transparency, which achieves competitive full-year prediction accuracy.
Abstract
Accurate coolant flow prediction is critical for active thermal management in high-performance computing (HPC) centers, yet it is inherently challenged by mixed-timescale dynamics and high-frequency workload surges. Existing deep learning methods often prioritize global accuracy on smoothed stationary trends, which may lead to phase delays during abrupt thermal transients. In addition, high-capacity architectures can introduce non-negligible computational overhead for latency-sensitive, resource-constrained edge controllers. To overcome these limitations, this study proposes a deployment-oriented edge-cloud collaboration (ECC) framework integrated with a transient-aware predictive architecture, named FS-Attention, designed to balance transient responsiveness, engineering deployability, and decision transparency. FS-Attention couples local feature synthesis, temporal-memory encoding, and attention-based temporal refinement to improve coolant-flow tracking under non-stationary operating conditions. Evaluations on the real-world Frontier supercomputer dataset show that the feature synthesis attention (FS-Attention) model achieves competitive full-year prediction accuracy, with a coefficient of determination (R²) of 0.8744 and a root mean square error (RMSE) of 0.0350. Under isolated critical thermal events (CTEs), FS-Attention obtains the lowest RMSE of 0.0757, slightly lower than the Temporal Fusion Transformer (TFT) and 5.61% lower than the standard Transformer. Platform-based profiling further shows an inference latency of 0.016 ms and a parameter size of 323.1 K, suggesting model-side compatibility with facility-side edge execution, while attention-shift analysis provides diagnostic evidence of temporally adaptive model behavior under dynamic thermal conditions.
Accurate job runtime prediction is essential for effective resource management in high-performance computing (HPC) systems. However, user-provided walltime estimates are often unreliable, and the operational costs of prediction errors are asymmetric.
This paper proposes a robust intelligent computing framework that con...
Hongyi Zhou· Poster Volume 0008 The 2026...· 0 citations
A double deep Q-Network-based proactive autoscaling approach (DDQN-Proactive) along with Resource Removal Strategy (RRS) that enhances decision-making by decoupling action selection from value evaluation, enabling more stable and adaptive scaling.
BOOSTEDSOSA is introduced, a dual-FPGA ML-assisted Scheduling architecture that integrates a Machine Learning predictor for expected processing times, with a novel temporal-aware training policy, enabling its use in existing HPC systems.
Adam H. Ross, Riccardo Revalor, Aryan Singh et al.· 0 citations
An LLM-driven, context-aware framework that integrates real-time system metrics, historical data, and task-specific importance levels for anomaly detection and prediction is proposed, enabling proactive intervention before critical operating conditions are reached.
Ioannis Tzitzios, A. Dimara, Georgiana Petridou et al.· Electronics· 0 citations
This paper proposes an edge-cloud collaborative physics-informed reinforcement learning framework for production data center HVAC control that integrates a physics-informed cold-start solution using Adaptive Particle Swarm Optimization, a three-time-scale edge–cloud architecture, and a constraint-aware safe projection...
Shichao Huang, Yi-Bing Zhou, Yuan Liu· Italian National Conference...· 0 citations
This work introduces a novel LLM-based predictive scheduling system designed to enhance operational efficiency while reducing the environmental impact of data centers, using an LLM to predict key metrics such as execution time and energy consumption from source code.
Hanzhao Wang, Jingxuan Wu, Yumeng Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.