Skip to content

Overwhelmed by Choice: Studying LLM Decision Making at Scale

Sep 2026 · 0 citations · 26 references
Computer Science

TL;DR

It is shown that strong small-option performance does not necessarily imply robust large-scale candidate comparison and that hierarchical partitioning and permutation-based inference improve accuracy by roughly 20 percentage points at $N=160$ on both HotpotQA and MIMIC.

Abstract

Multiple-choice and candidate-selection evaluations are widely used to assess LLM reasoning and decision-making, yet most benchmarks contain relatively small candidate sets. It remains unclear whether conclusions drawn from these settings remain valid as the candidate space scales. We systematically evaluate LLMs as the number of competing candidates increases and find substantial accuracy degradation across tasks, prompting strategies, and model scales. Controlled analyses show that standard long-context retrieval explanations cannot fully account for this degradation. Instead, we identify two systematic failure patterns. First, gold-margin collapse: the score gap between the correct answer and the strongest distractor progressively shrinks, driven primarily by weakening confidence in the correct answer. Second, earlier candidate preferences become increasingly difficult to overturn, with later candidates exerting progressively weaker influence on the final prediction. Motivated by these findings, we evaluate hierarchical partitioning and permutation-based inference, which improve accuracy by roughly 20 percentage points at $N=160$ on both HotpotQA and MIMIC. Overall, our results identify candidate-set scale as an important evaluation-protocol variable and show that strong small-option performance does not necessarily imply robust large-scale candidate comparison.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models

Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model must also use the ordinal decision scale supplied by the user faithfully. We analyze JEV~1.13 and three open KEV models...

Tian-Xiang Gao, Jin-Zhe Li, Zhiyuan Li et al. · 0 citations
Preprint Aug 2026

Accuracy and Order Sensitivity Diverge Under Label-Free Strategies

This paper test whether preventing a model from seeing option labels while committing to an answer removes positional influence and, in turn, improves performance, and evaluates two different strategies for mitigating bias.

Karleen Hanna, Feng Chen · 1 citation
#natural language process... Preprint Sep 2026

Knowing Is Not Choosing: What Explicit Verification Adds Beyond Generative Preference

Generating a correct answer does not mean that a language model will select it. We separate factual recall into three steps: generating a correct candidate, ranking the available candidates, and selecting the final answer. Pre-generation readouts predict factual recall and which questions sampling will cover across thr...

Yi-Long Li, Cheng-Po Yan, Aayan Arish et al. · 0 citations
Preprint Aug 2026

Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning

Funnel of Thoughts (FoT) is introduced, an inference-time method that preserves the full 32-trajectory voted accuracy while halving its attention FLOPs, a 28.8% reduction in full-model inference cost.

Chanhee Park, Sun Han, Jeongho Yoon et al. · 0 citations
#machine learning Preprint Sep 2026

OSCAR: Order-aware Scoring and Calibration for AI Rankings

Judge-specific sensitivity is useful for aggregating pairwise LLM evaluations, but its interpretation depends on which systematic presentation effects the ranking model includes. We introduce OSCAR, an order-aware framework for scoring and calibrating AI rankings, and study position as one such effect. In released judg...

You Liu, Yue Liu, Quanchao Lu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

LLMs are increasingly used as automated judges for model training and evaluation, yet individual judges exhibit systematic biases that undermine reliability. Much of prior work has studied biases in pairwise LLM-as-a-judge settings; in this paper, we focus on absolute scoring tasks, which mirror more realistic use case...

Gemma Zhang, Prachi Badarayani, Asmi Kumar et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.