Skip to content
Preprint

Ockhamareto: Pareto-Gated Segment-Level Credit Assignment for Concise Unit-Test Generation with Reinforcement Learning

Aug 2026 · 3 citations · 34 references
Computer Science

TL;DR

Ockhamareto is introduced, a single-shot GRPO framework for unit-test generation and selection, based on the principles of Ockham's Razor and Pareto Optimality, and it is shown that the knee point of the optimal trade-off between efficiency and effectiveness on the Pareto front is not correlated with obvious more easily computed proxy metrics, such as function size.

Abstract

We introduce \textbf{Ockhamareto}, a single-shot GRPO framework for unit-test generation and selection, based on the principles of \emph{Ockham's Razor} and \emph{Pareto Optimality}. Ockhamareto has two principal components: (i)~a \emph{Pareto-gated Bonus} that rewards only rollouts non-dominated in~(mutation, $-$\#tests) space, and (ii)~\emph{Token-level Segment Credit}, which attributes each test's marginal mutation kills back to the tokens of its unit-test block. On the \emph{UnLeakedTestBench~(ULT)}, Ockhamareto \emph{strictly Pareto-dominates} the strongest RL baseline~(\emph{MIST-RL}). Furthermore, it dominates on {\em each and all} optimization objectives, catching more bugs ($49.9\%$ vs $31.3\%$ mutation score at $N{=}5$), using \emph{fewer} tests ($2.60$ vs $4.67$ on average), thereby achieving $3.4\times$ the per-test trade-off improvement. The advantage is found in all four benchmarks~(\emph{HumanEval+}, \emph{MBPP+}, \emph{CodeContests}, \emph{TestGenEval-Lite}): Ockhamareto leads both mutation and coverage metrics on every one, always with the smallest suite. Ockhamareto also outperforms the state-of-the-art at all model scales, adding $+30$--$35$~pp mutation at 4B, 9B, and 27B model sizes. We also show that the knee point of the optimal trade-off between efficiency and effectiveness on the Pareto front is not correlated with obvious more easily computed proxy metrics, such as function size. This finding motivates the Pareto front computation; it is needed to identify this crucial engineering trade-off for each function under test.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

GLIDE: Generalized Layer-wise Intrinsic Distributional Evaluation for Heterogeneous LLM Agents

LLM agents require reliable step-level evaluation to compare candidate branches and allocate computation effectively. However, lightweight evaluation remains challenging. External verifiers introduce additional inference cost, while agent-produced confidence or self-evaluation scores can be miscalibrated, especially wh...

Wei Zhu, Yiming Wang, Rui Wang et al. · 0 citations
#machine learning Preprint Sep 2026

LLM Unlearning Evaluation with TRIAGE

TRIAGE can be applied alongside existing unlearning benchmarks to complement behavioral evaluation with a model-internal view of how unlearning reshapes the model's parameter space and affects retained knowledge.

Danial Ataee, Peter Triantafillou · 0 citations
Preprint Aug 2026

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents

A hierarchical credit assignment framework that retains RL's verifier-bounded ceiling while incorporating dense token-level signals from a privileged self-teacher is introduced, unlocking dense credit assignment without sacrificing the verifier-bounded ceiling.

Zechuan Wang, Siyuan Lu, Hongxuan Zhang et al. · 4 citations
#artificial intelligence Preprint Oct 2026

Instance-Dependent Regret for CMDPs with Step-Wise Constraints

We study online learning in episodic tabular constrained Markov decision processes with step-wise safety constraints. In such a setting, the constraints induce a safe subgraph that shapes the variance of cumulative rewards under feasible policies and, consequently, the difficulty of learning. Exploiting this structure,...

Qiang Zuo, F. Stradi, Le-Yang Xue et al. · 0 citations
Book Open access Sep 2026

Alignment + Accuracy: The Cascade Reward Representation for Preranking

Prerankers in large-scale recommender systems select candidates for a downstream ranker under strict latency constraints. In practice, teams combine accuracy metrics with alignment losses to train and evaluate prerankers, but what these quantities should target—and how to combine them—remains ad hoc. We derive the Casc...

Hedi Xia, Sheng-Lang Zhou, Ya-Li Bian et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.