Skip to content
Preprint

Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity

Aug 2026 · 0 citations · 29 references
Computer Science

TL;DR

This work probes models to answer two questions: do their activations encode whether a referent falls inside the knowledge boundary, and do they anticipate the specificity of the referent they are about to generate?

Abstract

When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to safer, more general claims. We frame this failure through a Gricean lens: a cooperative speaker who is uncertain about a referent retreats up the specificity hierarchy, trading informativeness for truthfulness. We ask whether LLMs have the ingredients to perform this retreat. Using a T-REx-based benchmark that varies entity familiarity and referent specificity, we probe models to answer two questions: (i) do their activations encode whether a referent falls inside the knowledge boundary, and (ii) do they anticipate the specificity of the referent they are about to generate? We find that the answer to both is yes, but the two signals are not reconciled in generation. Models overwhelmingly prefer specific referents even when the entity is unknown to them, and do so even when offered correct generic alternatives. The substrate for a Gricean retreat is present, but the policy that would act on it is not. We position our findings as a first step toward Gricean alignment, training or steering objectives that couple knowledge-boundary awareness to referent-specificity during generation.

View source

Similar papers

Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge

Factual question answering (QA) typically assumes a single canonical answer, obscuring whether large language models (LLMs) retain divergent accounts of long-tail facts. To address this gap, we introduce ElephantBench, a closed-book knowledge probe comprising 1,094 questions generated through an auditable graph-based p...

Zhuo-Shi Pan, Jun-Ru Lu, Yan-Fei Qian et al. · 0 citations
Open access Aug 2026

The Quantity‐Competence Presumption: Children Use Informativeness to Calibrate Trust and Infer Speakers' Knowledge

ABSTRACT Choosing whom to learn from is critical for human learners. A well‐established strategy to do so is to track informants' past accuracy. Yet this requires both a track record and independent knowledge of the truth—neither of which is available when encountering a new informant in an unfamiliar domain. This pape...

Cyann Bernard, Adeline Depierreux, Olivier Mascaro · 0 citations
#natural language process... Preprint Sep 2026

When Models Defer to Wrong Answers: A Robustness Audit of Source-Attributed Cues in Multiple-Choice QA

Language models often receive a question together with a claim about what another source answered. We audit whether such claims destabilize answers in multiple-choice question answering. For each item, we hold one wrong option fixed across misleading conditions and vary the cue template attached to it. We introduce \em...

Manikandan Ravikiran, Siddharth Vohra · 0 citations
Preprint Aug 2026

Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM Agents

The memory-clarification boundary is studied: whether interaction-derived information should be persisted, used only in the current context, re-verified, or clarified with the user, as well as across Claude and Qwen.

Bai-Chuan Li, Jun-Yi Yao, Zi-Hao Zheng · 3 citations
Preprint Aug 2026

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.

V. Rodionov, Shamil Assylbekov · 0 citations
#natural language process... Preprint Aug 2026

Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It

It is shown how fact-checking, a generally desirable behavior, can interfere with belief tracking in LLMs and how suppressing this attention at decoding time recovers accuracy only partially and only in some models, calling for future work on intervention methods.

Quang Minh Nguyen, Luis Frentzen Salim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.