Preprint
Aug 2026
Execution-grounded evaluation reveals hidden failures in language-model calculations for environmental science
AtmosCoder-Bench is introduced, an execution-grounded benchmark that makes the calculation process visible, and finds that multiple-choice formats inflate measured accuracy by at least 12 percentage points.
Maohao Ran, Chendong Ma, Yanting Zhang et al.
· 0 citations