Toward Trustworthy AI Software Evaluation: A Controlled Benchmark of Deep Learning Architectures for 24-h Photovoltaic Power Forecasting
A controlled and reproducible benchmarking framework that benchmarks fourteen forecasters under identical conditions, with explicit leakage safeguards, per-horizon reporting, and operationally meaningful peak diagnostics is presented, enabling claims of architectural superiority to be made trustworthy rather than merel...