Preprint
Aug 2026
TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs
A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.
V. Rodionov, Shamil Assylbekov
· 0 citations