Evaluating RL efficiency improvement methods including Synthetic Data Augmentation and Any-Generation Reward Optimization for Mathematical Reasoning on Countdown Tasks
This work investigates the effectiveness of synthetic data augmentation for Countdown arithmetic reasoning under both supervised and preference-based learning and develops a framework to evaluate whether synthetic supervision can improve reasoning performance without requiring additional human annotations.