Skip to content
Open access

APPLICATION OF LLM + ZERO-SHOT LARGE MODELS FOR FRUIT OBJECT DETECTION

Aug 2026 · INMATEH Agricultural Engineering · 0 citations · 11 references

TL;DR

A zero-shot annotation framework that integrates OWLv2, Google’s second-generation open-vocabulary vision model, with large language models to enable multilingual, natural language-driven fruit recognition in smart agriculture, providing scalable solutions for automated annotation, real-time monitoring, and large-scale data collection.

Abstract

Efficient and flexible agricultural image annotation is crucial for intelligent crop monitoring in smart agriculture, yet conventional detection models are limited by fixed class labels and require extensive manual annotations. This study presents a zero-shot annotation framework that integrates OWLv2, Google’s second-generation open-vocabulary vision model, with large language models (e.g., GPT-3.5, DeepSeek V1) to enable multilingual, natural language-driven fruit recognition in smart agriculture. A user-friendly interface was developed to support individual or batch image annotation with adjustable sensitivity to meet diverse field requirements. Experimental evaluations demonstrated the framework's strong generalizability and semantic understanding capabilities, allowing recognition of unseen fruit categories and attributes such as ripeness or color. The system significantly reduces annotation time and labor costs, while enhancing accessibility through natural language interaction. To ensure a robust evaluation of generalizability, a cross-domain protocol was employed using a novel dataset from 2025. Results showed that OWLv2 achieved an F1-score of 0.80 and an mAP of 0.8301, significantly outperforming the pre-trained YOLO11 (F1: 0.74, mAP: 0.60) in zero-shot scenarios. OWLv2 exhibited superior flexibility and required no task-specific dataset retraining, although its computational demands remain higher than lightweight models like YOLO11. Notably, while the LLM (DeepSeek) introduced a total one‑time API latency of 598.3 ms (called only once for processing multiple images). the actual core computational latency of OWLv2 was only 257.7 ms per image. Despite a total processing time of 891.9 ms (including visualization output), the framework demonstrates superior recall (0.9080) and semantic flexibility without retraining. These results verify the enormous application potential of OWLv2 and similar zero-shot models in agriculture, providing scalable solutions for automated annotation, real-time monitoring, and large-scale data collection.

Read PDF

Similar papers

Open access Jul 2026

A hierarchical prototype-graph with optimal-transport matching for few-shot rice disease recognition.

Accurate identification of rice diseases from field images is critical for crop health monitoring and sustainable agriculture, particularly in low-resource environments. However, most deep learning approaches depend on large-scale labeled datasets and pretrained backbones, limiting their applicability to rare or emergi...

M. D. Tanzimul Islam, Jobayar Alom, Masuduzzaman Niloy et al. · 0 citations
Open access Jul 2026

ITR: Iterative Transductive Refinement with a Graph-Reliability Gate for Zero-Shot Agricultural CLIP Classification

Contrastive Language–Image Pre-training (CLIP) has become the dominant paradigm for zero-shot visual recognition, classifying images directly from textual class descriptions with no task-specific labelled data. This label-free ability is a natural fit for agricultural plant-disease and weed recognition, where expert an...

Soo-Chang Lee, Jin Lee, Hoang Anh Le et al. · 0 citations
Open access Aug 2026

Fruit detection for small datasets via adjustable anchor boxes and transfer learning

Fruit detection is a crucial task in plant phenotyping but remains challenging due to limited training data, high variability in fruit appearances across different growth stages, and occlusions that hinder accurate detection. To address these issues, we propose an Adjustable Anchor Box Detection Network with Transfer L...

Dan Dai, Jun-Feng Gao, E. Sklar et al. · 0 citations
Review Open access Aug 2026

Deep Learning for In-Field Occlusion Handling and Real-Time Fruit Detection Under Dense Canopy Conditions

Robotic fruit harvesting in dense canopies remains challenging due to occlusion, variable illumination, and fruit-foliage similarity. This review synthesises recent deep learning-based detection systems, with particular focus on occlusion mitigation through multi-stage perception pipelines. The literature reveals that...

Abid Hayat, Shuvadeep Halder, Subham Ghosh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.