Skip to content
Preprint

Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors

Aug 2026 · 0 citations · 28 references
Computer Science

TL;DR

It is argued that the use of language to define the environment and rewards introduces unavoidable biases derived from the fact that the model is trained on word co-occurence, with implications for the reliability and robustness of LLM agents in real-world decision-making settings.

Abstract

Large language models (LLMs) are increasingly deployed as decision-making agents in settings that require sophisticated environmental exploration. However, existing work has raised questions about how LLMs actually balance exploration and exploitation. Unlike classical agents, LLM agents engage with tasks through natural language, exposing them to semantic information with no formal counterpart in the task structure. We introduce the semantic bandit, an extension of the multi-armed bandit setting that explicitly considers the textual labels assigned to actions, and use it to study how semantic priors --- inductive biases arising from associations between language and expected reward learned during pre-training, shape LLM exploration behaviour. We find that semantically informative action labels reduce exploration in favour of exploitation, improving performance when aligned with the reward structure and severely degrading it when misaligned. We further find that negative rewards trigger substantially more exploration than equivalent positive rewards, consistent with an expected-scale bias induced by reward conventions common in pre-training data. Overall, we argue that the use of language to define the environment and rewards introduces unavoidable biases derived from the fact that the model is trained on word co-occurence, with implications for the reliability and robustness of LLM agents in real-world decision-making settings.

View source

Similar papers

Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning

The findings relocate the bottleneck of commonsense reasoning failures from model competence to elicitation, and release SaliTrap as a testbed for this blind spot, to show that lightweight, inference-time prompting alone substantially closes the gap without any retraining.

Zheng Wu, Chen-Hao Xue, Shijie Zheng et al. · 0 citations
Preprint Aug 2026

BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?

The nature of test-time exploration in RLVR-trained LLMs is investigated by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence to delineate between entropy arising from stylistic variations and genuine inferential...

Soumadeep Saha, Krish Sharma, Akshay Chaturvedi et al. · 1 citation
Preprint Aug 2026

Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge

The Information Abundance Paradox is supported and it is suggested that scaling toward near-infinite context is not simply a matter of supplying more data, even when high-quality long-context data is abundant.

Arda Uzunouglu, Benjamin Van Durme, Daniel Khashabi · 2 citations
#artificial intelligence Review Aug 2026

A Survey on Rubric-Guided Reinforcement Learning for Language Models

A Bayesian framework that defines constitutions as prior distributions over evaluation criteria and rubrics as conditional instantiations is introduced, and a taxonomy of rubric-guided RL along the prior-posterior axis is presented, covering constitutional AI, instance-specific rubrics, process-level supervision, self-...

Zifei Shan, Fang-Ning Shao · 1 citation
Book

From Patterns to Semantic Anchoring

UCCT is framed as a Concept & Feasibility contribution offering a falsifiable pre­ dictive framework for when in-context control tips, with broader validation across multitask geom­ etry, RAG and adaptation interventions, and mechanistic localization left to follow-up studies.

Unknown authors · 0 citations
Preprint Jul 2026

ZenGen: Social Mind for LLMs

ZenGen, an integrated framework for measuring, internalizing, and grounding social intelligence, and Actio, a harness-controlled inference architecture that routes four typed supports into reasoning demonstrate the effectiveness of typed runtime support.

ZenGen Team, Xiang Ao, Jingping Bi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.