Skip to content

DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?

Sep 2026 · 0 citations
Computer Science

Abstract

Reliable automated research requires agents to vet data, verify analyses, and generate hypotheses grounded in trustworthy evidence, potentially reducing routine scientific workload while allowing scientists to focus on interpretation and discovery. Existing benchmarks often only assess analytical task completion or hypothesis generation separately rather than testing whether reliable evidence supports valid and novel claims. We introduce DISCERN (Data Integrity and Scientific Capability: Evidence, Reasoning, and Novelty), a controlled benchmark on real, publicly available datasets that evaluates three key levels of an automated research workflow. The first two levels test data integrity and analysis verification under confounds and tool traps, while the third tests hypothesis generation and revision under adversarial review, including counterfactual cases in which evidence consistent with real data and documented scientific phenomena conflicts with established expectations, motivating alternative explanations and testable hypotheses. Across 203 tasks, eight life-science tracks, and eight models, DISCERN shows that strong aggregate performance can mask level-specific weaknesses. Agents earn perfect scores in only 60.8% of Level 1, 34.2% of Level 2, and 0.6% of Level 3 evaluations, with penalties attributed to rejection of sound data, failure to carry recognized limitations into conclusions, and wide variation in hypothesis production. Cross-track rankings by token and code use are substantially more stable than rankings by evidence judgment, suggesting greater consistency in computational effort than in evidence-based reasoning. These profiles identify opportunities for supervised scientific assistance, but current agents do not yet demonstrate reliable autonomous analysis or discovery. Code and data: https://huggingface.co/datasets/discern-bench-anon/discern-benchmark

View source

Similar papers

Review Open access Sep 2026

The evaluation crisis in artificial intelligence‐driven computational biology: Pitfalls and solutions

Artificial intelligence (AI) is rapidly transforming computational biology. However, closer scrutiny reveals that many reported results do not hold under rigorous evaluation. This paper identifies prevalent, recurring issues in the field, including narrow datasets, hidden confounders that make supposedly independent...

T. F. Khang, Joshua W. K. Ho, Lin Hou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

Sci-MMR is introduced, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions, and it is found that current answer-centric benchmarks substantially overestimate the evidence-gro...

Jia-Qiang Li, Ya-Jie Yang, Zhi-Heng Xi et al. · 0 citations
Preprint Aug 2026

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.

V. Rodionov, Shamil Assylbekov · 0 citations
#artificial intelligence Preprint Oct 2026

From Scientific Observations to Mechanisms: Benchmarking Hypothesis Generation by AI Scientists

Data-driven mechanistic hypotheses are essential to scientific discovery because they explain how underlying processes produce observed phenomena. AI agents and AI scientists increasingly support scientific data analysis. However, their ability to turn empirical findings into mechanistic hypotheses remains insufficient...

Xia-Xun Xie, Qing-Qing Long, Meng Xiao et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.