Skip to content

Evaluating RL efficiency improvement methods including Synthetic Data Augmentation and Any-Generation Reward Optimization for Mathematical Reasoning on Countdown Tasks

· 0 citations · 12 references

TL;DR

This work investigates the effectiveness of synthetic data augmentation for Countdown arithmetic reasoning under both supervised and preference-based learning and develops a framework to evaluate whether synthetic supervision can improve reasoning performance without requiring additional human annotations.

View source

Similar papers

#large language models Open access Sep 2026

Toward scalable generative AI: efficient language model distillation via zero-shot rationales

This paper investigates an efficient approach for distilling Large Language Models (LLMs) into smaller, application-specific models using zero-shot Chain of Thought (CoT) rationale generation and Optimization by Prompting (OPRO). To address the challenges of deploying computationally intensive generative AI for narrow...

Lukas Vöge, Vincent Gurgul, Stefan Lessmann · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models

SOLID is proposed, a novel framework for self-improving OR language models without verified answers or external evaluators that improves solution accuracy for both general-purpose and OR-tuned models over outcome-only group-relative training.

Rui-Chen Zhu, Ming-Long Cao, Chen-Yu Zhou et al. · 0 citations
Conference Open access Sep 2026

Sprint or Delve: A Distribution-Aware Approach to Efficient Reasoning

The Powered Length Penalty (PLP) is proposed, an adaptive regularizer that penalizes redundancy in short sequences while gradually reducing penalties for longer sequences, preserving deep reasoning.

Ze-Hui Ling, De-Shu Chen, Hong-Wei Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

DE-Venus: A Data-Efficient RLVR Framework for Large Language Models

Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost of obtaining reliable targets at scale. Existing methods address sample selection, incomplete supervision, or noisy labels separately, ofte...

Shen-Zhi Yang, Guang-Cheng Zhu, Kai Tang et al. · 0 citations
#small language model Preprint Aug 2026

From Documents to Reasoning: A Validated Synthetic Data Pipeline and Semantic-Aware Fine-Tuning for Financial Numerical Reasoning

This work proposes a comprehensive pipeline for improving financial QA systems through high-quality synthetic data generation and fine-tuning of smaller language models (SLMs) using Quantized Low-Rank Adaptation (QLoRA).

Lokendra Birla, Milind Savagaonkar, Visnu Srinivasan et al. · 0 citations
#machine learning Preprint Aug 2026

Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs

This work proposes a synthetic, simulation-driven framework for studying knowledge insertion in LLMs, and introduces {\sc ParallelEvents}, a benchmark of fictional yet realistic future worlds that generates coherent event trajectories for controlled evaluation, avoiding contamination while preserving consistency.

Jonathan Zheng, Zi-Rui Shao, Alan Ritter et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.