Skip to content
Preprint

Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify

Aug 2026 · 0 citations · 26 references
Computer Science

TL;DR

This work formulate PRM stress testing as a quality-diversity search problem using MAP-Elites, retaining the most severe correctness-flipping edit in each behavior-space region while separating search coverage from exploit coverage, and characterize what such archives certify.

Abstract

Process reward models (PRMs) score intermediate reasoning steps and are widely used for search, ranking, and training, but optimization can exploit these learned proxies by increasing reward while turning correct reasoning into incorrect reasoning. We formulate PRM stress testing as a quality-diversity search problem using MAP-Elites, retaining the most severe correctness-flipping edit in each behavior-space region while separating search coverage from exploit coverage. We characterize what such archives certify: finite-cell repair bounds covered-cell tail risk and average residual severity but cannot bound the worst remaining cell from covered fraction alone; under Lipschitz post-repair loss and metric-cover auditing, the residual is bounded by archive fitting error plus the Lipschitz constant times the covering radius. A controlled landscape validates this certificate and the impossibility of any fraction-only worst-case guarantee. On real PRMs, the search reveals an aggregation-dependent vulnerability in Qwen2.5-Math-PRM-7B: padding yields 44 strict exploits with maximum gain 0.294 under mean pooling versus one exploit under minimum readout; a matched syntactic control isolates the mechanism, and an RLHFlow value-head model shows the same qualitative effect with maximum gain 0.005. A predeclared paired LoRA repair protocol reduces exploit rates from 0.148 to 0.037 to 0.074, lowers the worst attack from 0.333 to 0.177 to 0.212, improves ranking AUROC without degrading best-of-4 accuracy, attributes gains to adversarial fine-tuning rather than archive diversity, and is confirmed by independent unpaired replications (44 to 1, clean-split worst gain 0.0092, MATH-500 41 to 0, clean ranking 40/40).

View source

Similar papers

#artificial intelligence Preprint Sep 2026

MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents

This work presents MemRiskBench, a five-category risk taxonomy operationalized by deterministic trace grounded checks, instantiated as a 120-episode scripted benchmark with full trace logging and no LLM-as-judge on the pass/fail path, evaluated on five locally run quantized instruction-tuned models.

Jian-Hua Jiang, Dong-Bo Yuan, Weihua Li · 0 citations
#artificial intelligence Preprint Sep 2026

DepthBenchCAD: When Does Deeper Auditing Yield More Reliable Conclusions?

Generative CAD models are expected to remain behaviorally correct after parameter edits, so increasing the number of edit checks is often treated as a direct route to more reliable evaluation. Under a fixed budget, however, auditing each program more thoroughly reduces the number of tasks and independent generations th...

Hong-Ye Yang, Zhi-Hao Xie, Sheng-Jun Xiong et al. · 0 citations
#machine learning Preprint Aug 2026

PruneShift: A Framework for Evaluating Decision Reliability in Structured Pruning

PruneShift, an evaluation framework that separates broad predictive fidelity, fidelity near selector outputs, and the quality of the selected pruning decision, is introduced, showing why predictive fit, decision reliability, and pruning method quality require separate evidence.

Hao Ye, Gao-Peng Zhang · 0 citations
#artificial intelligence Preprint Sep 2026

Partition Scores Are Not System Scores: Deployment-Fidelity Gaps in Decomposed Algorithm Selection

Oracle-style quantities, including virtual best solvers, selected-portfolio VBS, virtual-best encodings, and best-in-family summaries, are widely reported as upper bounds on what a deployable selector could achieve. In decomposed algorithm selection, an analogous partition-level score grants an oracle choice of the bes...

Jiacheng Zhang, Yu Tang, Li Zhu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.