Evaluating Tree-Based Machine Learning Regression Models for Crop Yield Prediction Using Indian Agricultural Data
Abstract
Accurate crop yield prediction supports agricultural planning and resource allocation. This study evaluates three tree-based regression configurations Coarse, Medium and Fine Regression Trees for predicting district-level crop production in India. The analysis uses a single secondary tabular source, the open-access Crop Production in India dataset (Abhinand, 2019), restricted to the years 2000–2014, comprising over 250,000 state–district–year–season–crop records described by six predictors: State-Name, District-Name, Crop-Year, Season, Crop and Area. No meteorological, soil, irrigation or farm-management covariates are used. Non-finite entries in the Area and Production fields were converted to missing values and imputed by the median of the corresponding Crop × Season group; the response was then transformed as ln(Production + 1). The four high-cardinality categorical fields were encoded by out-of-fold target encoding with additive smoothing, the encoder being refitted inside each training fold to prevent leakage, and principal component analysis retaining 95% of variance was applied to the numeric block. Under 5-fold cross-validation the Fine Tree configuration achieved the best performance of the three arms, with R2 = 0.8402 ± 0.0029, RMSE = 0.3512 and MAE = 0.1948, all computed in the log-transformed response space. A corrected resampled paired t-test indicates that the difference from the Medium Tree configuration is unlikely to be attributable to fold-partition variability alone. The contribution is deliberately scoped as a controlled comparison of regression-tree complexity settings under a fixed, leakage-controlled pipeline on a national dataset; region-stratified models, ensemble baselines and residual diagnostics are outside the present scope and the resulting limitations are stated explicitly in Section 6.2.