Jul 2026· International Conference on the Theory of Information Retrieval· pp. 34-43· 3 citations· 51 references
Computer Science
TL;DR
The sandbox provides a search API that indexes large-scale public web corpora, namely ClueWeb22 and FineWeb, using a state-of-the-art dense retriever and approximate nearest neighbor search via DiskANN and achieves comparable latency to popular commercial APIs while ensuring stable document rankings across runs.
Abstract
Deep research systems represent an emerging class of agentic information retrieval methods that generate comprehensive and well-supported reports to complex queries, and/or answers to hard-to-locate factual questions. However, most existing systems rely on dynamic commercial search APIs, which pose reproducibility and transparency challenges, in addition to high costs. To address these limitations, we introduce DeepResearchGym as a free and open-source search sandbox for reproducible research on deep research systems. The sandbox provides a search API that indexes large-scale public web corpora, namely ClueWeb22 and FineWeb, using a state-of-the-art dense retriever and approximate nearest neighbor search via DiskANN. It achieves comparable latency to popular commercial APIs while ensuring stable document rankings across runs. We demonstrate the sandbox's utility through two use cases. For training, we synthesize queries grounded in the indexed corpora and show that search agents trained within the sandbox generalize to commercial search at inference time, enabling cost-effective reinforcement learning. For evaluation, we extend the Researchy Questions benchmark with LLM-as-a-judge metrics to measure alignment with users' information needs, retrieval faithfulness, and report quality. Evaluation results show that system rankings remain consistent when switching from commercial APIs to ours.
While autonomous agents have made significant strides in"deep research"by iteratively navigating the open web to synthesize information, real-world problem-solving is rarely confined to a single environment. Complex analytical tasks inherently require agents to weave together evidence from both ambiguous unstructured t...
Ruo-Fan Wu, Pei-Ran Xu, Xiao-Long Li et al.· 0 citations
This work introduces a verifiable benchmark of 500 deep research tasks spanning 31 topics and 10 major categories, with three query forms designed to probe complementary capabilities required for deep research.
Can Wang, Hao-Ran Chen, Hao Gao et al.· 0 citations
WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents that replaces static gold answer sets with task-specific judges that refetch cited pages and verify each record against its evidence, allowing evaluation of current and changing facts.
Vitaliy Polshkov, Marcin Pitera, Jeremy Yang et al.· 1 citation
Deep Research Pretraining (DRP), an offline framework that derives predictive navigation supervision from naturally occurring evidence structures, is introduced, an offline framework that derives predictive navigation supervision from naturally occurring evidence structures.
Jiang-Nan Zhou, Zhi-Yuan Fan, Xing Wu et al.· 2 citations
A taxonomy of core techniques, a layered system architecture and architectural paradigms, reviews representative implementations and applications, and highlights open challenges and future directions are provided.
Jinyan Cai· International journal of eng...· 0 citations
This work introduces S IEVE, a search–inspect–fetch strategy built around a Boolean Query Language (BQL), which improves accuracy with every tested ranker, and the accuracy–context advantage persists across retriever choices and agent backbones.
Shuai Wang, Hao-Dong Chen, Yu Yin et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.