StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models
This work proposes StrategyBench, which selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines evaluation metrics along two dimensions: strategy quality and downstream utility, and experiments show that explicit strategy utility differs substantially across task categories and depends on both strategy generation and execution conditions.
Jinghan Tan, Yuanzheng Wang, Lu Chen et al.
· 0 citations