Skip to content
Book Open access

DeepResearchGym: A Free, Transparent, and Reproducible Sandbox for Deep Research

Jul 2026 · International Conference on the Theory of Information Retrieval · pp. 34-43 · 3 citations · 51 references
Computer Science

TL;DR

The sandbox provides a search API that indexes large-scale public web corpora, namely ClueWeb22 and FineWeb, using a state-of-the-art dense retriever and approximate nearest neighbor search via DiskANN and achieves comparable latency to popular commercial APIs while ensuring stable document rankings across runs.

Abstract

Deep research systems represent an emerging class of agentic information retrieval methods that generate comprehensive and well-supported reports to complex queries, and/or answers to hard-to-locate factual questions. However, most existing systems rely on dynamic commercial search APIs, which pose reproducibility and transparency challenges, in addition to high costs. To address these limitations, we introduce DeepResearchGym as a free and open-source search sandbox for reproducible research on deep research systems. The sandbox provides a search API that indexes large-scale public web corpora, namely ClueWeb22 and FineWeb, using a state-of-the-art dense retriever and approximate nearest neighbor search via DiskANN. It achieves comparable latency to popular commercial APIs while ensuring stable document rankings across runs. We demonstrate the sandbox's utility through two use cases. For training, we synthesize queries grounded in the indexed corpora and show that search agents trained within the sandbox generalize to commercial search at inference time, enabling cost-effective reinforcement learning. For evaluation, we extend the Researchy Questions benchmark with LLM-as-a-judge metrics to measure alignment with users' information needs, retrieval faithfulness, and report quality. Evaluation results show that system rankings remain consistent when switching from commercial APIs to ours.

Read PDF

Similar papers

Benchmarking Hybrid Deep Research Across Database Querying and Web Search

While autonomous agents have made significant strides in"deep research"by iteratively navigating the open web to synthesize information, real-world problem-solving is rarely confined to a single environment. Complex analytical tasks inherently require agents to weave together evidence from both ambiguous unstructured t...

Ruo-Fan Wu, Pei-Ran Xu, Xiao-Long Li et al. · 0 citations
Review Aug 2026

WANDR: A Benchmark for Wide and Deep Research

WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents that replaces static gold answer sets with task-specific judges that refetch cited pages and verify each record against its evidence, allowing evaluation of current and changing facts.

Vitaliy Polshkov, Marcin Pitera, Jeremy Yang et al. · 1 citation
Preprint Aug 2026

Deep Research Pretraining via Predictive Navigation

Deep Research Pretraining (DRP), an offline framework that derives predictive navigation supervision from naturally occurring evidence structures, is introduced, an offline framework that derives predictive navigation supervision from naturally occurring evidence structures.

Jiang-Nan Zhou, Zhi-Yuan Fan, Xing Wu et al. · 2 citations
Review Open access 2026

DeepResearch: A Survey of LLM-based Research Agents

A taxonomy of core techniques, a layered system architecture and architectural paradigms, reviews representative implementations and applications, and highlights open challenges and future directions are provided.

Jinyan Cai · 0 citations

Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents

This work introduces S IEVE, a search–inspect–fetch strategy built around a Boolean Query Language (BQL), which improves accuracy with every tested ranker, and the accuracy–context advantage persists across retriever choices and agent backbones.

Shuai Wang, Hao-Dong Chen, Yu Yin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.