Skip to content
Open access

Toward Trustworthy AI Software Evaluation: A Controlled Benchmark of Deep Learning Architectures for 24-h Photovoltaic Power Forecasting

Jul 2026 · Computers · Vol 15, pp. 474 · 0 citations · 30 references

TL;DR

A controlled and reproducible benchmarking framework that benchmarks fourteen forecasters under identical conditions, with explicit leakage safeguards, per-horizon reporting, and operationally meaningful peak diagnostics is presented, enabling claims of architectural superiority to be made trustworthy rather than merely favourable.

Abstract

Accurate 24 h photovoltaic (PV) power forecasting is essential for day-ahead scheduling, storage operation, reserve planning, and market participation. However, published deep learning comparisons are often difficult to reproduce and interpret because they use inconsistent datasets, forecasting horizons, baselines, evaluation metrics, and leakage-control procedures. From a software engineering perspective, this limits the trustworthiness, comparability, and practical adoption of AI-based forecasting systems. This paper presents a controlled and reproducible benchmarking framework for evaluating AI-driven forecasting software. The framework is applied to nine deep learning architectures, three non-deep learning reference models, and two persistence baselines for hourly PV-power forecasting at a 350 kWp rooftop installation near Edinburgh, Scotland. All models were evaluated under a consistent experimental protocol, including the same chronological train–validation–test split, a 32-feature meteorological and solar-geometry input set, a 24-step forecasting horizon, capacity-normalised mean absolute error (NMAE), and Bayesian hyperparameter optimisation. The results show that TCN-LSTM achieved the best aggregate H24 performance with 7.22% NMAE, narrowly outperforming CPWformer-DEC at 7.28% and CT-PatchTST at 7.31%. LightGBM ranked fourth at 7.35% with fixed hyperparameters, outperforming six of the nine deep learning models. The top three models differed by only 0.09 percentage points, indicating that architectural superiority cannot be established reliably without significance testing and operational diagnostics. Per-horizon analysis showed that CT-PatchTST and S-Mamba performed best at the nearest forecast steps, whereas TCN-LSTM provided the most stable far-horizon profile. Peak-power diagnostics further revealed that aggregate NMAE can mask operational shortcomings, as Naive Persistence outperformed all deep learning models in high-output peak detection. The findings highlight the importance of reproducible benchmarking, leakage safeguards, horizon-aware evaluation, and operationally meaningful diagnostics in trustworthy AI software evaluation. The novelty of this work lies not in proposing a new architecture but in a controlled, reproducible framework that benchmarks fourteen forecasters under identical conditions, with explicit leakage safeguards, per-horizon reporting, and operationally meaningful peak diagnostics, enabling claims of architectural superiority to be made trustworthy rather than merely favourable. Architecture selection for PV forecasting should therefore consider not only aggregate accuracy but also reliability, interpretability of evaluation outcomes, and deployment-relevant performance behaviour.

Read PDF

Similar papers

Open access Aug 2026

Short-term PV power forecasting under real-world data constraints: a benchmark study of neural networks with uncertainty quantification

This study provides an in-depth comparative analysis of four state-of-the-art neural architectures, confirming that high-fidelity point forecasts and rigorously quantified uncertainty can be achieved simultaneously, providing a clear path toward more dependable PV dispatch, reserve allocation, and market participation.

Saloni Dhingra, G. Gruosso, G. Storti Gajani · 0 citations
Conference Open access 2026

Hybrid Stacking with Targeted Residual Learning for PV Forecasting

A hybrid PV forecasting framework that combines stacking ensemble learning with a targeted residual correction strategy, and demonstrates that analyzing error distribution and forecasting robustness provides valuable insights beyond conventional aggregate metrics, contributing to the development of more reliable photov...

Khawla Oufrit, A. Mouadili, M. Zazoui · 0 citations
#machine learning Preprint Aug 2026

An AI-Based Decision-Support Pipeline for Day-Ahead Photovoltaic Forecasting

A deployment-oriented environmental-AI pipeline for day-ahead hourly PV forecasting that corrects timestamp conventions, constructs leakage-safe solar-geometry and clearness-index features, adds short-term atmospheric context, and combines complementary predictors through validation-learned stacking is developed.

Fariba Dehghan, Sebastian Stein, V. Yazdanpanah et al. · 0 citations
Open access 2026

Solar Energy Production Forecasting With Hyperparameter-Optimized Time-Series Dense Encoder

This work introduces a fully automatic photovoltaic forecasting framework that combines the time-series dense encoder architecture with a genetic hyperparameter space explorer based on the non-dominated sorting genetic algorithm III to provide a promising approach for solar power forecasting.

Darius Peteleaza, Radu Sorostinean, A. Gellert et al. · 0 citations
Conference Aug 2026

Explainable Deep Learning for Sustainable Solar and Wind Power Forecasting

Solar PV and wind energy development in ongoing power systems is a means of achieving sustainability in power generation. Climatic uncertainty and renewables' variability, however, add to power forecasting volatility, which impacts energy management and reliability. Most existing deep learning models are hard to interp...

G. Kumaresan, N. C, S. R. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.