Skip to content

Optimizing information retrieval tasks with large language model for data enhancement

Aug 2026 · Journal of Nonlinear, Complex and Data Science · Vol 27, pp. 425 - 441 · 0 citations · 26 references

TL;DR

The goal is to optimize the normalized discounted cumulative gain (NDCG) metric, which measures the ranking quality of the retrieved documents, by integrating the fine-tuned LLM model with the chosen information retrieval system by incorporating the model’s outputs into ranking algorithms.

Abstract

Abstract The data science and artificial intelligence, optimizing information retrieval tasks has become crucial for extracting actionable insights from vast amounts of data. The problem is the need for precise query formulation to retrieve relevant data effectively, as LLMs can generate vast amounts of information that might include noise or irrelevant details. The objective of this study is to enhance the efficiency and accuracy of information retrieval tasks by leveraging large language models (LLMs) for data augmentation. Gathering a diverse dataset from various sources like online search engines, social media platforms, and online forums is crucial for meeting text and information needs effectively. The term frequency-inverse document frequency (TF-IDF) technique is applied to calculate the importance of each term in the dataset, allowing for the differentiation of significant words from common ones. This step is crucial in the data pre-processing phase to enhance the relevance and precision of information retrieval tasks. The goal is to optimize the normalized discounted cumulative gain (NDCG) metric, which measures the ranking quality of the retrieved documents. The framework integrates the fine-tuned LLM model with the chosen information retrieval system by incorporating the model’s outputs into ranking algorithms. The results show that the proposed method has the maximum accuracy, with an average accuracy of around 10 % when implemented using Python software. The future scope for optimizing information retrieval tasks with large language models (LLMs) for data enhancement is vast and promising.

View source

Similar papers

Conference Aug 2026

Evaluation of the BERT model for text semantic similarity

This study validates the effectiveness of the BERT model in semantic similarity calculation, providing more accurate technical support for related application scenarios, and laying the foundation for subsequent model optimization and lightweighting research.

Jia-Cheng Gao · 0 citations
Open access Aug 2026

Provisioning An Adaptive Model to Analyze Uncertainty and Large Language Patterns for Enhanced Document Re-Ranking

This work introduces a trust based adaptive reranking model- ATM (Adaptive Trust Model) that allocates computational resources according to file level uncertainty, instead of assigning a fixed number of reranker calls per query, which focuses computation only where ranking confidence is low.

Jenny Kalaiarasi.S · 0 citations
Book Open access Mar 2026

VILLA: Versatile Information Retrieval from Scientific Literature Using Large Language Models

This study designs a unique, open-ended SIE task of extracting mutations in a given virus that modify its interaction with the host, and develops a new, multi-step retrieval augmented generation (RAG) framework called VILLA for SIE.

Blessy Antony, Amartya Dutta, Sneha Aggarwal et al. · 0 citations
Review Open access Sep 2026

A Survey of Retrieval-Augmented Language Models for Knowledge-Intensive Text Applications

Retrieval-Augmented Language Models (RALMs) have emerged as an effective approach for addressing the limitations of conventional language models in knowledge-intensive text applications. These models combine external knowledge retrieval and language generation, enabling them to deliver more relevant, accurate, and cont...

Sachin Manekar · 0 citations
#natural language process... Preprint Aug 2026

Improving Information Extraction with Learned Queries

This paper shows that another part of the pipeline matters at least as much: the queries used to elicit information extraction, and introduces List of Questions (LoQ), which generates document-specific question sets, and FeedQ, a feedback-driven optimization method that iteratively refines questions against extraction...

Omar Sharif, S. Vosoughi, Nikhil Singh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.