Ockhamareto is introduced, a single-shot GRPO framework for unit-test generation and selection, based on the principles of Ockham's Razor and Pareto Optimality, and it is shown that the knee point of the optimal trade-off between efficiency and effectiveness on the Pareto front is not correlated with obvious more easily computed proxy metrics, such as function size.
Abstract
We introduce \textbf{Ockhamareto}, a single-shot GRPO framework for unit-test generation and selection, based on the principles of \emph{Ockham's Razor} and \emph{Pareto Optimality}. Ockhamareto has two principal components: (i)~a \emph{Pareto-gated Bonus} that rewards only rollouts non-dominated in~(mutation, $-$\#tests) space, and (ii)~\emph{Token-level Segment Credit}, which attributes each test's marginal mutation kills back to the tokens of its unit-test block. On the \emph{UnLeakedTestBench~(ULT)}, Ockhamareto \emph{strictly Pareto-dominates} the strongest RL baseline~(\emph{MIST-RL}). Furthermore, it dominates on {\em each and all} optimization objectives, catching more bugs ($49.9\%$ vs $31.3\%$ mutation score at $N{=}5$), using \emph{fewer} tests ($2.60$ vs $4.67$ on average), thereby achieving $3.4\times$ the per-test trade-off improvement. The advantage is found in all four benchmarks~(\emph{HumanEval+}, \emph{MBPP+}, \emph{CodeContests}, \emph{TestGenEval-Lite}): Ockhamareto leads both mutation and coverage metrics on every one, always with the smallest suite. Ockhamareto also outperforms the state-of-the-art at all model scales, adding $+30$--$35$~pp mutation at 4B, 9B, and 27B model sizes. We also show that the knee point of the optimal trade-off between efficiency and effectiveness on the Pareto front is not correlated with obvious more easily computed proxy metrics, such as function size. This finding motivates the Pareto front computation; it is needed to identify this crucial engineering trade-off for each function under test.
LLM agents require reliable step-level evaluation to compare candidate branches and allocate computation effectively. However, lightweight evaluation remains challenging. External verifiers introduce additional inference cost, while agent-produced confidence or self-evaluation scores can be miscalibrated, especially wh...
Wei Zhu, Yiming Wang, Rui Wang et al.· 0 citations
TRIAGE can be applied alongside existing unlearning benchmarks to complement behavioral evaluation with a model-internal view of how unlearning reshapes the model's parameter space and affects retained knowledge.
A hierarchical credit assignment framework that retains RL's verifier-bounded ceiling while incorporating dense token-level signals from a privileged self-teacher is introduced, unlocking dense credit assignment without sacrificing the verifier-bounded ceiling.
Zechuan Wang, Siyuan Lu, Hongxuan Zhang et al.· 4 citations
Experiments spanning mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks show that RISE outperforms RLVR-only training and on-policy self-distillation across all settings.
Yang Li, Semih Yavuz, Shafiq Joty· 3 citations· ⚡1
We study online learning in episodic tabular constrained Markov decision processes with step-wise safety constraints. In such a setting, the constraints induce a safe subgraph that shapes the variance of cumulative rewards under feasible policies and, consequently, the difficulty of learning. Exploiting this structure,...
Qiang Zuo, F. Stradi, Le-Yang Xue et al.· 0 citations
Prerankers in large-scale recommender systems select candidates for a downstream ranker under strict latency constraints. In practice, teams combine accuracy metrics with alignment losses to train and evaluate prerankers, but what these quantities should target—and how to combine them—remains ad hoc. We derive the Casc...
Hedi Xia, Sheng-Lang Zhou, Ya-Li Bian et al.· Proceedings of the 20th ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.