Evaluating RL efficiency improvement methods including Synthetic Data Augmentation and Any-Generation Reward Optimization for Mathematical Reasoning on Countdown Tasks
This work investigates the effectiveness of synthetic data augmentation for Countdown arithmetic reasoning under both supervised and preference-based learning and develops a framework to evaluate whether synthetic supervision can improve reasoning performance without requiring additional human annotations.
This paper investigates an efficient approach for distilling Large Language Models (LLMs) into smaller, application-specific models using zero-shot Chain of Thought (CoT) rationale generation and Optimization by Prompting (OPRO). To address the challenges of deploying computationally intensive generative AI for narrow...
Lukas Vöge, Vincent Gurgul, Stefan Lessmann· Management & Marketing· 0 citations
SOLID is proposed, a novel framework for self-improving OR language models without verified answers or external evaluators that improves solution accuracy for both general-purpose and OR-tuned models over outcome-only group-relative training.
Rui-Chen Zhu, Ming-Long Cao, Chen-Yu Zhou et al.· 0 citations
The Powered Length Penalty (PLP) is proposed, an adaptive regularizer that penalizes redundancy in short sequences while gradually reducing penalties for longer sequences, preserving deep reasoning.
Ze-Hui Ling, De-Shu Chen, Hong-Wei Zhang et al.· Proceedings of the Thirty-Fi...· 0 citations
Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost of obtaining reliable targets at scale. Existing methods address sample selection, incomplete supervision, or noisy labels separately, ofte...
Shen-Zhi Yang, Guang-Cheng Zhu, Kai Tang et al.· 0 citations
This work proposes a comprehensive pipeline for improving financial QA systems through high-quality synthetic data generation and fine-tuning of smaller language models (SLMs) using Quantized Low-Rank Adaptation (QLoRA).
Lokendra Birla, Milind Savagaonkar, Visnu Srinivasan et al.· 0 citations
This work proposes a synthetic, simulation-driven framework for studying knowledge insertion in LLMs, and introduces {\sc ParallelEvents}, a benchmark of fictional yet realistic future worlds that generates coherent event trajectories for controlled evaluation, avoiding contamination while preserving consistency.
Jonathan Zheng, Zi-Rui Shao, Alan Ritter et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.