Hybrid CNN-ViT framework for enhanced plant illness identification: accuracy and interpretability improvements
Abstract
The timely detection of plant health status is essential for achieving several benefits, such as increasing crop yield, reducing the use of toxic crop inputs, promoting healthier crops, and improving economic returns. A computer-aided plant status identification system enables plant health assessment using plant leaves, as they are the most prominent parts of the plant. Deep learning and convolutional neural networks (CNNs) are distinguished in the area of plant health identification. Still, CNNs cannot handle variable-sized images and struggle to extract features efficiently when the background is complex. To avoid these problems, this research proposes a hybrid model for identifying plant (groundnut) illness status using a CNN and a vision transformer (HCVT). CNNs are efficient at identifying local features, and vision transformers are efficient at identifying global features from leaf images; leveraging these strengths enhances the model’s efficacy. The HCVT model used the groundnut and rice datasets for experimentation. The HCVT model achieved 96.40% accuracy on the groundnut leaf image dataset and 99.86% on the rice dataset. The ablation analysis performed the necessity of each component in the HCVT model. The local interpretable model-agnostic explanations technique was used to understand the HCVT method’s functionality, and the results showed that the HCVT model outperformed current cutting-edge models in identifying plant disease and demonstrated its generalization potential. The HCVT model is affordable and widely accessible for analyzing plant leaf images. Utilizing the HCVT model offers a robust, easily accessible approach to diagnosing plant diseases by analyzing leaf images.