Skip to content

RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature

Aug 2026 · 0 citations · 107 references
Computer Science

TL;DR

RATIO (Retrieval Across Typed Ideation Operations), a large-scale benchmark in which relevance is defined by three operations which are name ideation moves, provides a scalable training and evaluation framework for retrieval components that support literature-grounded ideation, opening up new research avenues on scientific inspiration retrieval.

Abstract

Retrieved scientific literature can serve as inspiration for both human and AI scientists. Inspiration can take different forms: prior work may directly suggest how to address a problem, or surface directions at different levels of abstraction - zooming out to a more general view or zooming in to a concrete realization. We introduce RATIO (Retrieval Across Typed Ideation Operations), a large-scale benchmark in which relevance is defined by three operations which we name ideation moves: Address retrieves potential approaches for stated problems, Broaden retrieves more general formulations, and Specify retrieves concrete instantiations. RATIO is constructed from millions of full-text scientific papers across CS literature via a general recipe that extends discourse-marker distant supervision - previously used only for classification - to corpus-scale retrieval, combined with extensive LLM and human vetting. Experiments show that operation-specific fine-tuning substantially boosts retrievers but leaves much room for further improvements. RATIO provides a scalable training and evaluation framework for retrieval components that support literature-grounded ideation, opening up new research avenues on scientific inspiration retrieval.

View source

Similar papers

Preprint Aug 2026

MUSES: A Benchmark for Prospective Intellectual-Roots Retrieval

Scientific discovery depends on finding prior literature that shapes what comes next. Existing retrieval systems optimize for relevance and popularity, often favoring central papers over less familiar works that later prove generative. We introduce \textbf{MUSES}, a million-instance benchmark for prospective intellectu...

Rohan Pandey, Sunjae Kwon, Hong Yu · 0 citations
Conference Open access 2026

Decoupling retrieval quality from generative reasoning: A multi-dimensional benchmark of RAG architectures

Traditional search engine returns ranked lists for humans to interpret. Retrieval Augmented Generation pipelines go further, feeding retrieved context directly into large language models to synthesize knowledge rather than simply surface it. This study addresses a focused question: When the generative layer is held con...

Assmaa Moutaoukkil, Ali El Mezouary, A. Idarrou et al. · 0 citations
Preprint Aug 2026

SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG

We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19. Rather than proposing a new model, we evaluate a fixed scientific RAG pipeline across three corpus scales: 1,034 chunks (1K papers), 5,160 chunks (5K papers), and 15,480 chunks (15K pap...

Kaysarul Anas Apurba, Mahade Hasan, Rofiqul Alam Shehab et al. · 0 citations
Preprint Aug 2026

GEM: A Generative Embedding Model Bridging Reasoning and Retrieval

GEM is presented, a generative embedding model that augments retrieval through its own knowledge by explicitly reasoning about user intent and relevance criteria, and its generative nature allows test-time compute scaling via prompting to further enhance retrieval performance.

Zhili Shen, Craig Macdonald · 0 citations
Review Aug 2026

Can Retrievers Find the Same Paper from Different Aspects? A Multi-Aspect Full-Paper Scientific Retrieval Benchmark

An expert-validated benchmark for multi-aspect, full-paper retrieval that evaluates whether retrievers can consistently recover the same paper from queries targeting its motivation, method, and experimental findings, and proposes MAPLE-Synth, a retrieval-based in-context learning pipeline that leverages OpenReview disc...

Yiyang Wei, Fang Guo, Qiji Zhou et al. · 0 citations
Preprint Aug 2026

GRAFT: Graph-Distilled Generative Retrieval for Facet-Aware Scientific Literature Exploration

Scientific papers may relate by problem, method, result, or contribution, but document-level retrievers collapse these into a single similarity score without saying why they are related. Citation- and similarity-based retrieval alone also confines search to the neighbourhood of what is already known, whereas generative...

Italo Luis da Silva, Hanqi Yan, Yujing Wang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.