Preprint
Aug 2026
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
Measuring false alarms on human-verified-correct ProcessBench traces with the present task held byte-identical, it is found that a completed audit ->repair episode already in the model's context lowers false alarms in 15 of 15 model x wording combinations.
Parsa Mazaheri, Kasra Mazaheri
· 0 citations