Agri-VLG is proposed, a robust framework that synergizes few-shot annotated data with open-source knowledge for precise identification and significantly outperforms state-of-the-art methods, proving that graph-based refinement and external knowledge-guided retrieval are highly effective for detecting counterfeit products in few-shot scenarios.
Abstract
Deep learning has achieved remarkable success in computer vision, yet the scarcity of labeled data in specialized fields like counterfeit agricultural product identification remains a significant challenge. While humans can distinguish authentic goods from fakes using prior knowledge, machines often struggle with extreme data imbalance and domain gaps in task-specific datasets like TLU-States. To address this, we propose Agri-VLG, a robust framework that synergizes few-shot annotated data with open-source knowledge for precise identification. The architecture incorporates a Vision-Language Model (VLM) to retrieve high-quality web candidates from open data, effectively bridging the domain gap between generic imagery and agricultural tasks. To exploit the relationships between retrieved samples and target images, a Graph Attention Network (GAT) layer is employed to facilitate cross-image interaction and refine visual features. Our approach follows a two-stage process: Stage 1 involves finetuning to align representations, while Stage 2 focuses on classifier retraining to shift from imbalanced to balanced predictions. Extensive experiments on agricultural benchmarks demonstrate that Agri-VLG significantly outperforms state-of-the-art methods, proving that graph-based refinement and external knowledge-guided retrieval are highly effective for detecting counterfeit products in few-shot scenarios.
A zero-shot annotation framework that integrates OWLv2, Google’s second-generation open-vocabulary vision model, with large language models to enable multilingual, natural language-driven fruit recognition in smart agriculture, providing scalable solutions for automated annotation, real-time monitoring, and large-scale...
Ying-Dong Qin, Hao-Yu Song, Jing-Yi Li et al.· INMATEH Agricultural Enginee...· 0 citations
Vision-language models (VLMs) show promise for agricultural classification, but zero-shot performance on disease, pest, damage, quality, and species identification remains poor, and it is unclear whether this reflects weak visual features or a failure to connect them to domain knowledge. We build a benchmark of 116 dat...
Earl Ranario, Jared Smith, Lars Lundqvist et al.· 0 citations
While the landscape of closed-set oriented object detection has been revolutionized in recent years, extending this success to the instances of unseen categories, the core challenge of open-vocabulary object detection (OVD), remains a formidable hurdle. Prevailing methods often resort to a dedicated student detector to...
Yan-Qing Yao, Xiang Yuan, Gong Cheng· IEEE Transactions on Geoscie...· 0 citations
A visual–semantic multimodal fusion framework is developed, incorporating a Two-Stage Modality Fusion mechanism that projects semantic and visual features into a unified feature space to optimize cross-modal feature interaction.
Hou-Kui Zhou, Xin-Zhang Li, Yuan Nie et al.· AgriEngineering· 0 citations
The Robust and Fine-grained training framework for CLIP-based vision-language models (RoFLIP) is proposed, enhancing both the robustness and granularity of vision-language alignment and underscore RoFLIP’s compositional reasoning and generalization abilities.
Yiwei Sun, Chuan-Bin Liu, Shancheng Fang et al.· International Journal of Com...· 0 citations
Results on the UC Merced (UCM) and NWPU benchmarks indicate that SE-CLIP significantly outperforms existing semi-supervised approaches and provides a viable solution for adapting VLMs to the remote sensing domain with minimal human intervention.
M. L. Mekhalfi, M. M. Al Rahhal, Y. Bazi et al.· IEEE Geoscience and Remote S...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.