Regime-Aware Model Selection for Internal Combustion Engine Performance Prediction Using Machine Learning
Abstract
Artificial Neural Networks (ANNs) and tree-based ensembles such as Random Forest (RF) and Extreme Gradient Boosting (XGBoost) are the two dominant families used for tabular regression in engineering informatics. The empirical literature, however, reports contradictory verdicts, on some datasets ANNs dominate, while on others RF and XGBoost are markedly superior, and no unifying account exists of the conditions under which each family is preferable. This paper addresses that gap directly. Using a curated dataset of 1,232 internal combustion engines, we formulate the prediction of rated power output as a supervised regression problem and derive, from first principles, the inductive biases of the three model classes. We show mathematically that the two families differ in the geometry of the function class they approximate ANNs build globally smooth compositions of ridge functions, whereas tree ensembles build axis aligned, locally constant partitions and that this distinction predicts where each should excel. Empirically, the three models achieve near-identical global accuracy (test R-squared of 0.957, 0.952 and 0.954 for ANN, RF and XGBoost respectively), yet a regime-stratified analysis reveals systematic, statistically meaningful reversals of rank across operating conditions, ANNs win in turbocharged, petrol and high-power regimes, while tree ensembles win in naturally aspirated, diesel, small displacement and low to mid-power regimes. We then train a logistic meta-model that predicts, from engine descriptors alone, whether a tree ensemble will outperform the ANN for a given instance, achieving 66% routing accuracy and identifying displacement, compression ratio and forced induction as the decisive covariates. An oracle router that selects the locally best model reduces mean absolute error from roughly 11 hp to 7.2 hp, a 33% reduction, quantifying the value of condition aware model selection. The study contributes a reproducible, mathematically grounded methodology for explaining and exploiting the ANN-versus tree performance gap rather than merely reporting it.