Explainable AI for transportation process management: symbolic regression of freight train speeds using Kolmogorov–Arnold networks (KANs)
Sergey E. EliseevNikolay A. DavydovMikhail P. NoskovSergey V. Eroshenko
Aug 2026· Railway Sciences· 0 citations· 13 references
TL;DR
Beyond the first railway application of KAN, the study presents a complete framework for explainable capacity management: it shows how a single KAN model can be distilled into a human-readable, operationally meaningful equation that quantifies nonlinear, asymptotic and interaction effects – a capability not offered by post hoc explainable artificial intelligence methods.
Abstract
This study addresses the black-box problem in railway predictive analytics by applying Kolmogorov–Arnold networks (KANs) to freight train speed prediction. The aim is to obtain not only an accurate forecast, but also an analytical expression that can support transparent operational and managerial decisions.
A comparative analysis of five approaches is conducted: linear regression, gradient boosting (CatBoost), multi-layer perceptron (MLP), symbolic regression using genetic programming (gplearn) and the proposed KAN architecture. The dataset comprises 55,508 observations with 23 operational and infrastructure factors. Three cross-validation strategies (random, time-based and track-stratified) and the Friedman test are employed. Symbolic formulas are extracted from linear regression, KAN and gplearn.
The trained KAN model (before symbolic extraction) achieves RMSE = 4.12–4.25, comparable to CatBoost (3.87–3.98) and superior to MLP (4.53–4.68). An analytical formula containing exponential, logarithmic and polynomial terms is successfully extracted from KAN. The RMSE of the KAN-derived formula is 13% lower than that of linear regression and 14% lower than that of gplearn on the full 23-feature set. The results confirm that KAN can capture complex dependencies while remaining analytically interpretable. The study presents a complete framework for explainable capacity management, offering a human-readable equation that quantifies nonlinear, asymptotic and interaction effects.
Beyond the first railway application of KAN, the study presents a complete framework for explainable capacity management: it shows how a single KAN model can be distilled into a human-readable, operationally meaningful equation that quantifies nonlinear, asymptotic and interaction effects – a capability not offered by post hoc explainable artificial intelligence methods. The resulting expressions can directly inform timetable optimisation, infrastructure investment appraisal and operational scenario analysis.
Accurate freight demand forecasting is essential for improving logistics planning and supporting decision-making in e-commerce supply chains. This study compares the performance of two forecasting approaches—Seasonal Autoregressive Integrated Moving Average (SARIMA) and Extreme Gradient Boosting (XGBoost)—for predicting daily freight demand measured by transported weight. The research followed the CRISP-DM methodology using the Brazilian Olist public e-commerce dataset. After data preprocessing, exploratory analysis, stationarity testing, and feature engineering, multiple SARIMA and XGBoost models were developed and evaluated using chronological train-test splitting, cross-validation, and Mean Absolute Percentage Error (MAPE). The SARIMA models incorporated seasonal differencing and Box-Cox transformations, whereas the XGBoost models included calendar-based variables, moving averages, and moving standard deviations. The results demonstrate that feature engineering substantially improved predictive performance. The best XGBoost model achieved a MAPE of 3%, considerably outperforming the best SARIMA model, whose predictive accuracy remained limited despite data transformations. These findings indicate that machine learning techniques combined with temporal feature engineering provide superior freight demand forecasts for e-commerce logistics. The proposed approach offers a practical decision-support tool for transportation planning, resource allocation, and operational efficiency while providing a reproducible computational workflow through publicly available source code and processed data.
Eduardo Modesto de Melo, F. Piurcosky· Revista Mythos· 0 citations
An integrated prediction-and-visualisation pipeline that transforms complex data distributions into actionable visual analytics, such as interpretable station-to-station demand heatmaps via interactive GIS Folium layers is implemented, providing an operationally robust framework to support smart-city transportation management and build more sustainable urban transit systems.
Berna Çalışkan· Journal of Data Analytics an...· 0 citations
The fast pace of development of e-commerce has elevated the timely delivery as a characteristic element of customer satisfaction and logistics performance. However, there are still delays in shipment because of uncontrollable factors like traffic, weather, operational bottlenecks and network inefficiencies. To overcome this issue, this paper derives a Machine Learning Model of Shipment Delay Prediction to combine refined shipment data, operational time and contextual logistics data to predict the probability of delay at an early phase. Based on the previous studies of real-time delay prediction, proactive risk assessment, as well as ML-based logistics optimization, the suggested framework will integrate feature engineering, supervised learning models (Random Forest, XGBoost, CatBoost, Logistic Regression), and a multi-stage prediction process. This methodology is focusing on interpretability, prediction on each shipment processing step, and scalability to the logistic operations. The experimental findings indicate that the gradient-boosting models are rather consistent in terms of their performance (high ROC-AUC scores and higher recall in the delay class). This study adds a useful and empirical methodology, which can be adopted by logistics teams to predict disruptions, make sound-informed routing, and enhance service reliability.
Pranjul Vishwari, Rajiv N Thakker, Sumit Verma et al.· International Conference on...· 0 citations
Maritime accidents such as capsizing, storm-induced roll resonance, collisions and groundings continue to occur in Bangladesh’s inland and coastal waterways. While these events are usually linked to overloading, weather conditions, and maintenance issues, another important hydrodynamic factor - the Added Mass Coefficient (AMC) is rarely examined. Traditional methods for calculating AMC are too slow for use in real operations. In this study, we explored how different machine learning (ML) models, including Random Forest (RF), Neural Networks (NN), Support Vector Regression (SVR), Linear Regression (LR), and Long Short-Term Memory (LSTM) networks can predict AMC values from basic vessel parameters. Our results show that AMC can be considered not only as a design variable but also as an operational safety parameter. By predicting AMC in advance the models provide a way to support safety actions such as adjusting heading, controlling load distribution and reducing risks in shallow-water navigation. We also suggest that future work should combine real-time data with hybrid approaches to strengthen the reliability of predictions.
M. Kabir, Al-Amin Hossain Seam, Z. I. Awal· Engineering· 0 citations
This study proposes an explainable AI approach for predicting localized congestion and vessel fuel demand to support sustainable port logistics, using Automatic Identification System (AIS) data. Raw AIS data were pre-processed and processed into operational features at the vessel level to create two target variables of interest related to logistics, including Congestion Index (CI), and Relative Fuel Index (RFI). Three ensemble machine learning algorithms, namely Random Forest (RF), XGBoost, and LightGBM, were implemented and evaluated using the coefficient of determination (R2) and root mean square error (RMSE). The Random Forest model demonstrated the best performance in congestion prediction with an R² of 0.8328 and an RMSE of 2.3200, followed by XGBoost with an R² of 0.8151 and LightGBM with an R² of 0.8110. XGBoost showed the best performance in predicting fuel, with an R2 of 0.9984 and an RMSE of 2.8164, followed by LightGBM (R2 = 0.9971) and Random Forest (R2 = 0.9960). SHAP explainability revealed that historical traffic density was the most important factor contributing to congestion. At the same time, the speed of vessels was the most important factor affecting the fuel demand. The proposed framework supports optimized vessel scheduling, shorter port waiting times, fuel savings, and carbon emissions reduction, which help build intelligent and sustainable green maritime logistics.
Hoang Phuong Nguyen, Chi Linh Duong, Q. Nguyen et al.· JOIV: International Journal...· 0 citations
Although data analytics and artificial intelligence are increasingly reshaping logistics decision-making, limited empirical evidence explains how organisations transform these technologies into tangible intelligence-driven performance, particularly in emerging economies. This study aims to develop and empirically examine the concept of logistics intelligence performance (LIP), defined as a multidimensional organisational capability that captures a firm’s capacity to transform logistics data and AI tools into timely, accurate and forward-looking decisions.
The study is based on survey data collected from 2,752 Moroccan firms. A quantitative approach is adopted, combining confirmatory factor analysis (CFA) to validate the measurement model of LIP with descriptive statistical analysis and a comparative machine learning framework. Multiple predictive models are implemented and evaluated, including linear regression, decision trees, ensemble methods (Random Forest, Gradient Boosting, XGBoost), support vector machines, k-nearest neighbors and artificial neural networks, to compare their ability to explain variations in LIP.
The CFA confirms the validity and reliability of the LIP construct, showing strong model fit and convergent validity. Machine learning models outperform linear regression in predicting LIP, with artificial neural networks achieving the highest accuracy (R² up to 0.74). Results indicate that data quality, analytical capabilities and alignment between analytics and decision-making are the strongest determinants of LIP. AI adoption alone has limited impact without complementary organisational capabilities. Findings also reveal significant non-linear relationships, threshold effects, and capability complementarities shaping the development of logistics intelligence across firms.
This study contributes to the literature by conceptualising and empirically validating LIP as a distinct organisational capability. By integrating organisational capability theory with machine learning methods in the context of an emerging economy, the research moves beyond technology adoption perspectives and provides new insights into how firms convert data and AI investments into intelligence-driven logistics performance.
Taoufiq El Moussaoui, Alaâ Eddine El Moussaoui· The International Journal of...· 0 citations