Skip to content

SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores

Sep 2026 · 0 citations · 11 references
Computer Science

TL;DR

SciRIGOR is introduced, an evaluation framework and benchmark comprising 100 cases from scientific articles across six domains and 17 subfields and scores complete claim-support paths while localizing the earliest unsupported relation.

Abstract

Scientific coding agents produce interdependent code, results, figures, and claims, yet evaluating final outputs alone does not establish whether their conclusions are scientifically supported. We formulate evidence-grounded multimodal scientific analysis, requiring agents to produce executable analyses and claims supported by results and visualizations from the same run. We introduce SciRIGOR, an evaluation framework and benchmark comprising 100 cases from scientific articles across six domains and 17 subfields. The framework reconstructs typed evidence graphs, separates artifact fidelity from relational validity, and scores complete claim-support paths while localizing the earliest unsupported relation. Source-grounded alternative paths accommodate scientifically equivalent analyses and visualizations. We evaluate 11 agent/model configurations. On full-benchmark runs, claims agree with faithful and unfaithful results at nearly identical rates (91.8% versus 91.0%). Yet no system exceeds 62.6% on the soft evidence-chain score or 18.0% strict whole-chain success. These findings show that internal coherence does not establish scientific correctness: evaluation must verify support along the complete data-to-claim path.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Can LLMs Discover Scientific Laws in Real and Parallel Worlds?

It is found that predictive fit can diverge from scientific validity, memorization shapes whether models reproduce or move beyond published formulas, and the best-of-N study reveals a selection bottleneck.

Yi-Ming Huang, Zi-Chen Liu, Junxia Cui et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

Sci-MMR is introduced, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions, and it is found that current answer-centric benchmarks substantially overestimate the evidence-gro...

Jia-Qiang Li, Ya-Jie Yang, Zhi-Heng Xi et al. · 0 citations
#natural language process... Preprint Sep 2026

EviScope: Paired Counterfactual Evidence Diagnostics for Faithful and Efficient Grounded Language Models

Grounded language-model systems are often evaluated by final answer accuracy, yet a correct answer can be unsupported, drawn from the wrong source, or produced when evidence is insufficient or contradictory. We introduce EviScope, a paired counterfactual benchmark that holds the question fixed while adding, removing, d...

Suryadeep Singh Deswal · 0 citations
Preprint Sep 2026

SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

This work introduces SciDocBench, a workflow-centered benchmark targeting capability diagnosis with verifiable training-data construction for scientific-document assistants, and SciDocIR, a structured representation of scientific document objects, layout and cross-reference relations, and provenance.

Shenxi Wu, Yu-Hong Liu, Hao-Song Zhang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.