Attention-Fused ConvNeXt and MobileViT Framework with FastSAM Lesion Extraction for Multi-Class Crop Disease Detection
Abstract
The problem of crop disease recognition from leaf images is a difficult multi-class visual classification problem due to disease evidence appearing at various spatial scales and the possibility of image content other than disease symptom including healthy tissue, background information, and visually similar symptoms. This paper presents an attention-fused dual-branch framework that combines Fast Segment Anything (FastSAM) based lesion extraction, ConvNeXt based full-leaf representation learning, and MobileViT based lesion-focused representation learning. The experimental study adopted a controlled subset of 4000 images from the 20k+ Multi-Class Crop Disease Images collection and applied class normalization and removed one small class, leaving 41 classes. There were 2777 training images, 588 validation images, and 635 testing images in the final split. The candidate masks were obtained by FastSAM, and then filtered and ranked with lesion-oriented criteria, then the selected region was cropped and refined by performing Lab-space contrast enhancement, robust chromatic deviation, Otsu thresholding and morphological operations. The original leaf was processed by ConvNeXt and the refined ROI processed by MobileViT. They were projected into a common 256-dimensional feature space and aggregated using a learnable two-branch gate that includes entropy regularization. The Fusion Model got 93.07% accuracy and 85.27% macro F1 on the common test set, while ConvNeXt got 89.76% accuracy and 81.55% macro F1, and MobileViT got 72.76% accuracy and 68.07% macro F1 on the common test set. The Fusion Model also achieved a one-versus-rest macro-average AUC of 0.9978. Grouped confusion matrices indicate that the fusion approach reduces off-diagonal errors at the crop-group level. A prototype of a Streamlit deployment was also created to showcase the following features: upload of images, lesion preprocessing, prediction, reporting of confidence, inspection of branches, and image visualization with Grad-CAM. The results support the use of a combination of global and lesion-specific representations rather than a single visual pathway.