Large language models (LLMs) are increasingly used as graders, verifiers, and process auditors, but most mathematical evaluations still emphasize final-answer accuracy. This can obscure whether a model can verify a non-canonical but valid solution trace. We introduce a controlled linear-equation benchmark for evaluating LLMs in the evaluator role. Each instance asks the model to judge final-answer correctness, step-level trace correctness, and the first incorrect step. Our evaluation of state-of-the-art open LLMs reveals a significant robustness gap: models that accurately evaluate canonical solutions often fail when presented with perturbed but logically equivalent variants. Across GPT-OSS 20B, Qwen3-14B, and Phi-4-Reasoning, base models perform well on canonical traces but degrade substantially on perturbed traces, especially for error localization. On valid perturbed traces, base-model false-rejection rates reach 75.6-85.3%, showing strong sensitivity to canonical solution form. Supervised fine-tuning, distillation, and test-time compute improve robustness in some settings, but gains are model dependent and can trade off against canonical performance. The results show that reliable process-level verification remains challenging, and evaluator robustness should be measured separately from solver accuracy, even in a simple algebraic domain with exact ground truth.
Crop modelling is essential for agricultural water management but often relies on simplified water balance routines that limit representation of soil moisture dynamics. To address this limitation, we developed a coupled model that integrates the 1‐D Richards equation, solved using a finite difference method into the FAO AquaCrop. The coupled model was calibrated and validated using soil moisture, canopy cover, above‐ground biomass and seed cotton yield data from field experiments in the southeastern United States. Compared with hourly field measurements of soil moisture, AquaCrop–Richards achieved an average root mean square error (RMSE) of 0.023 m
3
m
−3
across three soil depths over the growing season. Model performance for canopy cover, biomass and yield resulted in RMSE values of 12.18%, 1.77 t ha
−1
and 0.96 t ha
−1
, respectively, against observations. Under fully irrigated conditions, both models produced statistically indistinguishable yield estimates. However, under rainfed conditions, AquaCrop simulated 15.5% higher yields than AquaCrop–Richards. Analysis showed that AquaCrop produced rapid stepwise drainage, resulting in root‐zone water content 33%–37% lower than the coupled model. This reduced soil moisture triggered earlier water stress which led to yield overestimation. These results indicate that AquaCrop‐Richards improves soil moisture representation and is robust under water‐limited conditions.
Krishna Panthi, Vidya Samadi, Carlos Toxtli· Irrigation and Drainage· 0 citations