Skip to content
Open access

Synergistic effects of spectral preprocessing and machine learning algorithms for nitrogen estimation in tomato using hyperspectral spectral data

Sep 2026 · Scientific Reports · 0 citations

Abstract

Accurate estimation of leaf nitrogen content is essential for optimizing fertilization management and improving crop productivity. This study proposes a nondestructive hyperspectral imaging framework combined with advanced machine learning algorithms to classify nitrogen levels in tomato leaves (Solanum lycopersicum L., Royal variety). A total of 300 hyperspectral samples were collected from plants subjected to three nitrogen treatments (N-30%, N-60%, and N-90%) under controlled conditions. Reflectance data (400–1000 nm) were calibrated and preprocessed using Standard Normal Variate (SNV), Multiplicative Scatter Correction (MSC), and Savitzky–Golay (SG) filtering. Both unsupervised and supervised algorithms were systematically evaluated. Clustering results demonstrated that preprocessing substantially influenced class separability. The MSC + GMM combination yielded the best unsupervised results, with the lowest Davies–Bouldin index (0.443) and the highest silhouette coefficient (0.975), indicating improved cluster compactness and separation. In supervised learning, Neural Networks (NN) consistently outperformed the other models, achieving 100% accuracy and an ROC–AUC of 1.00 under SNV, MSC, and SG preprocessing. Gradient Boosting (GB) and Support Vector Machine (SVM) also demonstrated strong predictive capability, whereas conventional models, including LDA, LR, KNN, and NB, showed moderate improvements following preprocessing. Statistical analysis confirmed the significant effect of spectral preprocessing on both clustering and classification outcomes ( p  < 0.05). Overall, integrating hyperspectral imaging with appropriate preprocessing and nonlinear machine learning models provided strong classification performance under the controlled experimental conditions. However, the exceptionally high performance observed for some models should be interpreted cautiously because of the relatively small dataset and experimentally controlled nitrogen treatments. Moreover, the absence of an independent external validation dataset limits the assessment of generalizability across growing conditions, cultivars, and field environments. Therefore, independent validation using larger and more diverse datasets is required before broader application of the proposed framework in practical precision agriculture.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.