Preprint
Aug 2026
SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models
State-of-the-art models are far more proficient in scientific coding than SciCode has suggested---the bottleneck was not model capability, but the quality of the evaluation instrument.
Sihan Hu, Lyuhan Huang, Youjin Deng et al.
· 0 citations