Preprint
Jul 2026
PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents
It is argued that evaluations of scientific agents should report not only accuracy, but also item-level retention, output-access sensitivity, trajectory failures, and where the computation chain breaks.
Ke Zhang, Sahchit Chundur, M. J. Qomi et al.
· 0 citations