Skip to content

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

Jul 2026 · arXiv.org · Vol abs/2607.07690 · 0 citations · 29 references
Computer Science

TL;DR

Agon is introduced, which makes two competing models each other's graders, which faces a progressively stronger rival, which single-model RL cannot provide.

Abstract

Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answer. On hard problems this trains models to write more rather than to think better, since the trace itself is never graded and no label for good thinking exists. We introduce Agon, which makes two competing models each other's graders. Both attempt the same problem; in alternating roles, one drafts a solution and the other reads it while solving, and each is rewarded for out-solving the other. To win, a model must out-reason a rival that has seen its work, so reasoning is judged implicitly during training, with no process labels and no reward model. Because both models are optimized, each faces a progressively stronger rival, which single-model RL cannot provide. The two need only be comparably strong and behaviorally different. At inference the pair deploys as it trains, a two-stage cascade in which one model drafts and the other answers after reading the draft. On the hard split of DeepMath with Qwen3, this doubles GRPO's pass@1, roughly eight times the gain of an untrained Mixture-of-Agents pass over the same base. The ordering replicates on competitive-programming code and across model families (Qwen3.5, Gemma 4). For now the models talk in text; the next step is to let them reason together in latent space.

View source

Similar papers

#machine learning Preprint Sep 2026

Cliff: Learning Process Rewards from the First Mistake

Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout, is proposed and established as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.

Pei-Xuan Han, Runnan Wang, Ketan Ramaneti et al. · 1 citation
Preprint Aug 2026

Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation

A reasoning model is built that adaptively chooses how much to reason for each problem, and the brief modes end up more accurate than \textsc{Long}, which shows that the router sorts problems by difficulty rather than at random.

Gijs Kassenaar, Zhao Yang, Vincent François-Lavet · 1 citation
#machine learning Preprint Sep 2026

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

Reinforcement learning from verifiable rewards (RLVR) drives chain-of-thought reasoning in large language models, yet its binary outcome reward cannot distinguish among correct trajectories. Existing dense reward alternatives, from surface heuristics to process reward models, either ignore the expert solutions already...

Le-Qi Zheng, Jin-Bo Su, Fang Niu et al. · 2 citations
Preprint Aug 2026

BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?

The nature of test-time exploration in RLVR-trained LLMs is investigated by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence to delineate between entropy arising from stylistic variations and genuine inferential...

Soumadeep Saha, Krish Sharma, Akshay Chaturvedi et al. · 1 citation
Preprint Aug 2026

IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents

Isolated Bilateral Reinforcement Learning (IB-RL), in which the two roles coevolve through joint rollouts while each role optimizes its own reward through fully independent advantages, action masks, and update paths, produces policies that generalize more effectively to unseen counterparts.

Senhao Wang, Chenghao Cai, Hai-Tao Hu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.