Overwhelmed by Choice: Studying LLM Decision Making at Scale
It is shown that strong small-option performance does not necessarily imply robust large-scale candidate comparison and that hierarchical partitioning and permutation-based inference improve accuracy by roughly 20 percentage points at $N=160$ on both HotpotQA and MIMIC.