Skip to content
Review Open access

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Nov 2025 · bioRxiv · 0 citations · 31 references
Biology

Abstract

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI’s ChatGPT, Google’s Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

Read PDF

Similar papers

Review Open access Aug 2026

Evaluating large language model performance in US FDA regulatory science

This study compares the performance of three LLMs in extracting, analyzing, and synthesizing regulatory and clinical information from FDA drug reviews, guidance for the industry, and drug labels as accessed through their standard user interfaces, using antibiotics approved for complicated urinary tract infections betwe...

Khulud Bukhari, R. Rodriguez-Monguio, B. Lopez-Bermudez et al. · 0 citations
Open access Sep 2026

LLM-enabled Natural History Study Analysis to Support Rare Disease Research

Background: Rare diseases affect an estimated 300 million people worldwide, yet the research needed to guide diagnosis and treatment is often fragmented across multiple unstructured literature sources. Natural history studies (NHS) are a key source of this evidence, but manually extracting structured information from N...

Kevin Li, E. Sid, Qian Zhu · 0 citations
Review Sep 2026

Harnessing the Power of Large Language Models for Drug Discovery: A Systematic Review of Current Applications and Future Directions.

INTRODUCTION The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the poten...

Pranay Wal, M. Shukla, Kritika Sachan et al. · 0 citations
Open access Aug 2026

Intelligent Framework for Adverse Drug Event Identification Using Large Language Models and Retrieval-Augmented Generation: Development and Evaluation Study

Synergizing a curated domain-specific knowledge base with LLMs via a RAG architecture is an effective strategy for accurately identifying ADEs in unstructured Chinese clinical notes, providing a foundational open-source benchmark and a robust technical framework to advance pharmacovigilance, drug safety research, and c...

Junlong Ma, Xue-Hong Wu, Zeying Feng et al. · 0 citations
Open access Sep 2026

Performance, Failures, and Oversight of a Large Language Model Agent for Clinical Data Analysis: Evaluation Study

Abstract Background Large language model (LLM) agents capable of generating and executing statistical code from natural language may broaden access to clinical data analysis, yet which pipeline stages they perform reliably and which require expert oversight remain poorly defined. Objective This study aimed to evaluate...

Yi-Lan Wu, D. J. Fu, Yu-Kun Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.