Skip to content

ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction

Sep 2026 · 0 citations · 16 references
Computer Science

TL;DR

This work presents ToolGate, which turns repeated answer checking and difficulty screening into an auditable process while leaving domain design and final review to experts.

Abstract

Scientific benchmarks are commonly built by domain experts who write tasks and cross-check one another's work, or who adapt existing material from textbooks, published papers, and online resources. These routes can produce strong evaluations, but they require substantial per-item labor. Language models can reduce this repeated work by proposing candidates quickly. The remaining problem is acceptance. We target scientific questions whose answers require computations with specialist software rather than unaided reasoning alone. A candidate is invalid if its script fails or returns a different answer, or trivial if a model answers it without the software. We present ToolGate, which treats every generated item as a proposal and keeps it only if three gates pass. First, an executable solution script must reproduce the proposed answer when run with the scientific software. Second, randomized no-tool screening rejects candidates that models can already solve from the prompt alone. Third, a tool-using agent must solve each survivor within a fixed time limit. We instantiate ToolGate in FEniCSx with 500 generation attempts. The local-verification gate retains 478 candidates. For final reporting, we rescreen this pool after generation: two randomized no-tool screens exclude 222 from the reported pool, and direct GPT-5.5 API calls at medium reasoning (the API default) exclude another 121. Of the remaining 135, a GPT-5.5 Codex CLI agent with access to FEniCSx solves 130; exact deduplication leaves 128 unique protocol survivors. ToolGate turns repeated answer checking and difficulty screening into an auditable process while leaving domain design and final review to experts.

View source

Similar papers

Review Open access Sep 2026

Automatic Model Selection Based on Task Complexity

An Adaptive Intelligence Router is set out that chooses among deterministic programs, retrieval and service interfaces, small language models, more capable language models, more capable language models, and accountable human review.

L. Santhanam · 0 citations
#artificial intelligence Preprint Sep 2026

TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science

Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretical computer science (TCS) connects algorithm design to explicit guarantees and fundamental limits, providing a setting for evaluating whether models can justify co...

Chu-Tong Yang, Xi-Yuan Zhang, Yu Huang et al. · 0 citations
Preprint Aug 2026

MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models

MemToC, a controlled benchmark for post-tool-return arbitration with executable tools, is introduced and an asymmetric success criterion is applied: correct-answer retention must improve without a detected reduction in correct-tool following.

A. Varlamov, R. Zinnatullin, Elisei Rykov et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Adaptive-GEPA: Make Your Harness Fit Heterogeneous Requests

Reflective optimizers such as GEPA improve language model prompts from execution traces and evaluator feedback; full-program extensions can also rewrite tools and control flow. In practice, a user hands the same endpoint heterogeneous requests whose effective solutions require different tools, reasoning modes, and cont...

Tian-Yu Chen, Yasi Zhang, Rui-Yi Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

An Empirical Evaluation of Cost-Efficient Large Language Models on Algorithmic Programming Tasks

This study empirically evaluates whether cost-efficient Large Language Models (LLMs) can be trusted to generate enterprise code to a written specification. Three models (Gemini Flash 3, GPT-5.4 mini and Claude Haiku 4.5) were asked to solve 992 algorithmic problems as Java Spring Boot service methods conforming to a ma...

Chandimal Adikari, Nandika Herath · 0 citations
#artificial intelligence Preprint Sep 2026

E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

Coding agents powered by large language models (LLMs) are evolving from making localized code changes to developing complete software repositories. However, evaluating repository-scale generation remains challenging: tasks must demand system-level reasoning while ensuring that all evaluated behaviors are precisely spec...

Hantian Ding, Chloe Bi, Jia-Cheng Zhu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.