Skip to content

Short-term wind power forecasting with lag-based feature engineering: a turbine-level machine learning comparison

Aug 2026 · Neural computing & applications (Print) · Vol 38 · 0 citations · 46 references

TL;DR

A turbine-level forecasting framework with a 10-minute temporal resolution, integrating unified multi-turbine data, lag- and rolling-based feature engineering, trigonometric time encoding, and machine learning models including Ridge, Lasso, K-Nearest Neighbors, Random Forest, Extra Trees, and XGBoost is developed, offering a reproducible solution for real-world short-term wind power forecasting.

View source

Similar papers

Aug 2026

SCADA-based feature analysis and Bayesian-optimized machine learning for short-term turbine-level wind power forecasting

Most wind power forecasting studies rely primarily on historical turbine output time series, while environmental and operational parameters remain insufficiently explored. This paper investigates short-term turbine-level wind power forecasting using SCADA data collected from a wind farm in Vietnam, incorporating both environmental and operational variables. A filter-based feature selection method is first applied to identify significant predictors based on p-values. The ARIMAX model is adopted as a statistical baseline. In addition, several machine learning models, including Gaussian process regression, support vector machine, random forest, least-squares boosting, and feedforward neural network, are implemented. Bayesian optimization is employed for hyperparameter tuning. Forecasting performance is evaluated using MAE and RMSE, which are further combined into a weighted objective function. Results indicate that, within the investigated one-month SCADA dataset, the optimized machine-learning models achieved lower MAE and RMSE values than the traditional ARIMAX baseline. These findings demonstrate the effectiveness of the proposed framework for short-term turbine-level forecasting under the investigated operating conditions, while broader seasonal generalization requires validation using longer-term SCADA datasets.

Ho-Tuan Le · 0 citations
Conference Aug 2026

Short-Term Wind Power Forecasting Using Transformer-XGBoost Hybrid Model

Accurate wind forecasting is critical to ensure stable and efficient integration of renewable energy resources in modern power systems. However, the inherent variability and non-stationarity of wind pose a significant forecasting problem for modern power system operators to ensure power system stability. A new hybrid forecasting model incorporating a Transformer Encoder and XGBoost regressor is proposed in this paper to enhance one-hour-ahead short-term forecasting of wind power generation. The model employs Transformer Encoders to exploit complex temporal relationships from sequences of 24-hour past wind data, learning daily and periodic trends, ramps, and hidden dynamics. The generated temporal feature vectors are then combined with meteorological variables such as wind speed and direction at various altitudes and passed on to the XGBoost model to obtain regression forecasts. The proposed approach has been validated on real-world datasets obtained from commercial wind energy farms. Experimental results demonstrate that the hybrid model outperforms standalone Transformer and XGBoost models by achieving accurate values for mean squared error, mean absolute error and coefficient of determination. In addition, model has tested for long-term forecasting as well. Although the model is effective for short-term forecasting, due to increased uncertainty and error propagation, its performance degrades with extended forecasting horizons.

Heshan Senapriya, Sakun Rasilka, D. P. Wadduwage · 0 citations
Conference Aug 2026

A Comparative Study of Ensemble Tree-Based Models for Short-Term Electricity Load Forecasting

Short-term load forecasting (STLF) is an essential task for reliable power system operation, economic dispatch, reserve scheduling, and grid planning. This study aims to provide an operationally realistic and interpretable comparison of five ensemble tree-based machine learning (ML) models for national electricity demand forecasting using the publicly available Panama Short-Term Electricity Load Forecasting dataset. Gradient Boosting Regressor (GBR), XGBoost, LightGBM, CatBoost, and Random Forest are evaluated using 14 predefined walk-forward train–test splits that emulate the weekly forecasting protocol of Panama’s national grid operator. A common feature set consisting of lagged demand variables, a four-week moving average, temporal indicators, calendar variables, and Tocumen temperature is used for all models. A seasonal naive baseline, statistical significance testing, COVID-period split analysis, and feature importance comparison are also included. CatBoost achieved the best average performance with an RMSE of 55.52 MWh and MAPE of 3.80%, outperforming the seasonal naive baseline, which obtained an RMSE of 78.61 MWh. However, Wilcoxon-Holm testing showed that the narrow RMSE differences among the ensemble models were not statistically significant at the 5% level. Feature importance analysis confirmed that the four-week moving average is a dominant predictor for most models. The results show that ensemble tree-based models provide accurate, robust, and interpretable STLF performance under an operationally realistic evaluation protocol.

Timur Lale · 0 citations
#artificial intelligence Preprint Sep 2026

Predicting Wind Turbine Power Using Machine Learning and Weather Forecasts

Offshore wind turbines are widely used to generate renewable energy, but their maintenance can result in decreased efficiency due to forced shutdowns. Accurate wind turbine power predictions can identify periods of low power that would be ideal for scheduling maintenance. However, the effects of data volume, feature selection, and data preprocessing on the performance of such power prediction models have not been thoroughly studied. Besides, current models have limited transferability between different wind turbines. Therefore, this study developed a baseline Linear Regression for performance comparison with a more complex Artificial Neural Network model to predict the power output of a wind turbine, using weather conditions only to enhance applicability. A range of data preprocessing techniques were studied, and models were trained on one month and one year of data to determine the effects of data preprocessing and volume on model performance. Feature selection was explored using a Random Forest Regressor. The best results from the different models showed that the Artificial Neural Network models provided the highest accuracy, with an R2 score of 0.98 and a low Mean Absolute Error of 194, when compared with the baseline model (R2 score of 0.94 and Mean Absolute Error of 441). The model performance is comparable to the range of results in past studies, with the advantage that the proposed method leverages a separate weather dataset from a nearby weather station, enabling future applications for similar wind turbines in different locations. The Artificial Neural Network model was then used to identify 4-h periods of low power predictions over 2 months (simulating application for future periods), providing power output savings of approximately 2000 kW for each maintenance event.

Khivishta Boodhoo, Isaac Triguero, Josh Plumbly et al. · 0 citations
Open access 2026

Enhanced hybrid residual learning framework for robust wind turbine power prediction using machine learning

Wind turbine power generation for any nation is always a priority since its impact on power grids and imbalance the sustainability parameters. This research work captures the variation in wind resource variability by implementing advanced machine learning models for short-term power forecasting using SCADA data. Here, a fresh framework was developed to focus on parameters such as aerodynamic behavior, temporal patterns, and overall regime. This information was fed to the model to analyze actual and theoretical power output. With the strong literature study few models were shortlisted such as Extreme Gradient Boosting (XGBoost), Deep Neural Networks (DNN), and a novel hybrid stacking ensemble to fulfill the criteria. From the generated data hybrid model was the top choice as its value for RMSE came down to 0.0103, while it moves as high as 0.9992 in case of R². With the need to optimize the present work, multiple algorithms were selected and merged to get the desired output so that accuracy can be maintained without compromising on the grid stability and performance. The models were so selected that the outcome can give lower reliance on carbon-intensive power. This work shows a better data driven based farmwork which only addresses energy forecasting parameters by relating to renewable penetration and sustainable energy systems.

Mohammad Y. Mhawiash, B. Khassawneh, Kamal Alieyan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.