Skip to content
Book Open access

VILLA: Versatile Information Retrieval from Scientific Literature Using Large Language Models

Mar 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 10396-10407 · 0 citations · 65 references
Computer Science

TL;DR

This study designs a unique, open-ended SIE task of extracting mutations in a given virus that modify its interaction with the host, and develops a new, multi-step retrieval augmented generation (RAG) framework called VILLA for SIE.

Abstract

The lack of high-quality ground truth datasets to train machine learning (ML) models impedes the potential of artificial intelligence for science research. Scientific information extraction (SIE) from the literature using LLMs is emerging as a powerful approach to automate the creation of these datasets. However, existing LLM-based approaches and benchmarking studies for SIE focus on broad topics such as biomedicine and chemistry, are limited to choice-based tasks, and focus on extracting information from short and well-formatted text. The potential of SIE methods in complex, open-ended tasks is considerably under-explored. In this study, we use a domain that has been virtually ignored in SIE, namely virology, to address these research gaps. We design a unique, open-ended SIE task of extracting mutations in a given virus that modify its interaction with the host. We develop a new, multi-step retrieval augmented generation (RAG) framework called VILLA for SIE. In parallel, we curate a novel dataset of 629 mutations in ten influenza A virus proteins obtained from 293 scientific publications to serve as ground truth for the mutation extraction task. We demonstrate VILLA's superior performance using a novel and comprehensive evaluation and comparison with vanilla RAG and other state-of-the art RAG- and agent-based tools for SIE. Finally, we evaluate the generalizability of our proposed method using another unique dataset of mutations in hepatitis E virus.

Read PDF

Similar papers

Review Open access Sep 2026

Hybrid clustering framework for large-scale scientific literature structuring

This study addresses the challenge of organizing and interpreting large collections of scientific literature, focusing on a rapidly growing and heterogeneous AI research domain derived from object-detection-with-small-data literature. First, based on the set of keywords using Clarivate Analytics queries, over 18,500 re...

Dmytro Teplov, Artur Radzivil, Andrej Bugajev · 0 citations
Aug 2026

Optimizing information retrieval tasks with large language model for data enhancement

The goal is to optimize the normalized discounted cumulative gain (NDCG) metric, which measures the ranking quality of the retrieved documents, by integrating the fine-tuned LLM model with the chosen information retrieval system by incorporating the model’s outputs into ranking algorithms.

C. Vaidya, Amudhavel Jayavel, Pradeep Kumar Mishra et al. · 0 citations
Review Open access Aug 2026

Applications of Natural Language Processing: A Comprehensive Study

A comprehensive review of the evolution of NLP from traditional rule-based approaches to modern transformer models including BERT and GPT demonstrates that NLP continues to transform intelligent systems and is expected to play an increasingly significant role in the development of next-generation AI technologies.

P. Kalaiselvi · 0 citations
Review Sep 2026

Abstract A099: Creating an interconnected pancreatic cancer research knowledge system using a Large Language Model and wiki publishing platform

The authors sought to utilize Large Language Models (LLMs) to interconnect previously separate areas of pancreatic cancer research into a dynamic research tool for the research community. The historical weaknesses of LLMs include hallucinations, dependence on unverified internet sources, and lack of determ...

Bernard A. Kroll, Joshua P. Raff, E. Fisher · 0 citations
#large language models Open access Sep 2026

Research on Text Information Extraction and Imbalanced Classification Methods for Enterprise Profiling

A comprehensive natural language processing (NLP) pipeline for extracting key information and identifying industries, which outperforms traditional machine learning baselines and single deep learning models, offering more reliable recognition for minority classes.

Xin-Yi Xu · 0 citations

Approaches for Extracting Research Infrastructure Information from Text

It is shown that the usage of CRF weights in BERT-based architectures achieves noteworthy improvements in the overall NER task by approximately 12 %, and that in few-shot learning set-ups the effectiveness of CRF weights is much higher in smaller training sets.

Georgios Cheirmpos, Seyed Amin Tabatabaei, Evangelos Kanoulas et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.