Hyperspectral monitoring of wheat gluten index based on multi-model ensemble and adaptive learning.
Wheat gluten index is an important indicator of food quality inspection. Although hyperspectral technology offers non-destructive and rapid detection potential, its application faces challenges including high-dimensional data redundancy, multicollinearity among adjacent wavelengths, and model overfitting risks in small-sample scenarios. This study proposes a rapid prediction method for wheat gluten index integrating adaptive spectral preprocessing with machine learning. Based on 89 wheat flour samples, adaptive wavelet denoising (average SNR improvement: 11.92 dB), successive projections algorithm (SPA) feature selection, and variance inflation factor (VIF) collinearity diagnosis reduced 2001 spectral features to 5 core variables (587 nm, 436nm_square, 1241 nm, 1080nm_square, 940nm_square). Ten algorithms including PLSR, SVR, XGBoost, and CatBoost were systematically compared under 70% training/30% test split with 10-fold cross-validation. CatBoost achieved optimal performance (R2 = 0.8552, RMSE = 7.8602), surpassing PLSR (R2 = 0.7090) and XGBoost (R2 = 0.7923) by 20.6% and 8.0% respectively. Notably, the well-tuned single CatBoost model outperformed stacking ensemble methods (R2 = 0.7803), providing empirical evidence for model selection in small-sample spectral analysis. SHAP interpretability analysis identified critical spectral bands corresponding to protein and starch absorption characteristics, offering guidance for portable equipment optimization and mechanistic understanding of spectral-quality relationships.