Skip to content
Open access

A retrieval-augmented and distribution-imbalance-aware contrastive framework for low-resource agricultural pest and disease named entity recognition with large language models

Aug 2026 · Journal of King Saud University: Computer and Information Sciences · Vol 38 · 0 citations · 40 references

TL;DR

A parameter-efficient large model adaptation framework that integrates retrieval-augmented semantic conditioning modeling with distribution-imbalance-aware contrastive representation learning for agricultural pest and disease text entity recognition is proposed, offering an efficient and scalable pathway for adapting large models to agricultural knowledge graph construction and intelligent decision-making in smart agriculture.

Abstract

Amid the rapid advancement of smart agriculture, a substantial volume of unstructured knowledge embedded in agricultural pest and disease texts urgently necessitates structured representation through high-quality named entity recognition. However, Chinese agricultural domain entities exhibit pronounced long-tail distributions, coupled with scarce annotated samples and semantic boundaries heavily reliant on implicit domain-specific knowledge structures. These challenges lead to performance degradation and insufficient generalization capabilities of traditional sequence labeling models and general pre-trained models under low-resource scenarios. To address these issues, this paper proposes a parameter-efficient large model adaptation framework that integrates retrieval-augmented semantic conditioning modeling with distribution-imbalance-aware contrastive representation learning for agricultural pest and disease text entity recognition. The proposed method reframes agricultural named entity recognition as a structured prediction problem conditioned on semantic neighborhood variables. By constructing query-relevant semantic neighborhoods and organizing the retrieved examples as contextual demonstrations, the proposed method provides retrieval-conditioned references for more stable entity boundary determination. At the representation learning level, a category-aware contrastive optimization mechanism is introduced, prioritizing the construction of semantically similar hard negative samples to reshape the geometric structure of the embedding space and mitigate frequency-dominated optimization biases induced by long-tail distributions. For model adaptation, a low-rank parameter-efficient fine-tuning strategy is employed to enable controlled transfer of large language models to the agricultural domain, reducing training costs while preserving general semantic capabilities. Extensive experiments under multi-gradient low-resource settings are conducted on two Chinese agricultural pest and disease datasets, AgCNER and CropDiseaseNER. Experimental results demonstrate that the proposed framework significantly outperforms both traditional sequence labeling methods and conventional large model fine-tuning strategies across varying data scales. Specifically, compared with the BERT-BiLSTM baseline, the proposed framework achieves F1-score improvements of 6.75 percentage points in the AgCNER-3k low-resource scenario and 15.30 percentage points in the CropDiseaseNER-0.4k extreme low-resource scenario. These findings indicate that retrieval-augmented semantic conditioning modeling and distribution-imbalance-aware contrastive representation optimization collaboratively mitigate structural instability arising from low resources and long-tail distributions, offering an efficient and scalable pathway for adapting large models to agricultural knowledge graph construction and intelligent decision-making in smart agriculture.

Read PDF

Similar papers

Conference 2026

Ambiguity-Aware Keyword-Enhanced Label-Aware Semantic Fusion for Text Classification

An Ambiguity-Aware Semantic Fusion Framework (AAK-LASFNet) for robust text classification and demonstrates that the proposed framework consistently outperforms strong baselines in terms of Accuracy, highlighting its effectiveness in alleviating semantic ambiguity in text classification.

Peijun Xie · 0 citations
#artificial intelligence Preprint Sep 2026

URA-NER: A Unified Retrieval-Augmented Framework with Retrieval Alignment and Uncertainty Reduction for Low-Resource NER

A novel unified retrieval-augmented framework, URA-NER, including three key components: Progressive Granularity Retrieval, Model-aware Representation Enhancement, and Reason-aware Knowledge Verification is proposed, including three key components: Progressive Granularity Retrieval, Model-aware Representation Enhancemen...

Jing-Yu Wang, Shijie Wu, Fu-Sheng Jin · 0 citations
Aug 2026

STaR: a soft-labeling and triplet-aware retriever for efficient retrieval-augmented QA

This study proposes STaR, a novel retriever fine-tuning framework that integrates BM25 similarity graph-based soft labeling with a triplet similarity learning strategy based on Sentence-BERT (SBERT), and introduces a triplet-aware SBERT training architecture that explicitly models relative semantic distances between qu...

Jiali Jiang, Chih-Yung Chang, Youxi Li et al. · 0 citations
#natural language process... Preprint Sep 2026

Domain-Adaptive Pretraining Enhances Water Treatment Semantic Representation for Large-Scale Structured Literature Mining

Water treatment research is expanding rapidly, but much of the knowledge acquired from this research remains scattered across unstructured literature. The field still lacks a dedicated language model that can efficiently capture water treatment-specific domain semantics for large-scale literature mining. Here, we addre...

Mu-Di Zhai, Rui-Hong Qiu, Qing-Yun Zeng et al. · 0 citations
Review Open access Aug 2026

Improving Indian Address Parsing in Data-Scarce Environments Using Chunk-Based Retrieval-Augmented Transformer Models

Accurate parsing of unstructured Indian addresses remains challenging due to the linguistic variability and the limited annotated data. While transformer-based models achieve near-perfect performance on synthetic datasets, their generalization to real-world inputs is limited, with F1-scores degrading substantially for...

Nikhil S. Dhavase, Navendu Chaudhary · 0 citations
Open access Aug 2026

A Method for Dynamic Expansion of Domain Dictionaries and Weakly Labeled Named Entity Recognition for Low-Resource Specialized Corpora

In order to solve the problems of domain entity lack and high noise, little human supervision in low resource specialized corpora, this work adopts the method of multi-feature fusion of dynamic domain dictionary expansion. The method includes dictionary matching, semantic confidence filtering and confidence-weighted na...

Y.-J. Ma · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.