Skip to content

Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models

Jul 2026 · arXiv.org · Vol abs/2607.12739 · 0 citations · 26 references
Computer Science

TL;DR

ESFP, a behavioral benchmark that treats the contrast between externally attributed and self-attributed prompts as the fundamental unit of measurement, is introduced, a behavioral benchmark that measures a model's propensity to adapt its epistemic stance under changing attribution conditions, rather than a general competence measure.

Abstract

A language model may be asked either what experts believe about a contested claim or what it believes about the claim itself. A trustworthy conversational agent should distinguish these two requests and respond in different epistemic registers: neutral attribution in the first case and stance expression in the second. Whether such a shift occurs-and whether it occurs coherently-is not directly assessed by existing benchmarks for accuracy, instruction following, or safety. We introduce ESFP, a behavioral benchmark that treats the contrast between externally attributed and self-attributed prompts as the fundamental unit of measurement. ESFP consists of 104 carefully controlled items spanning six epistemic categories and five phrasing templates, and evaluates model responses along four complementary dimensions: lexical self-attribution, representation-level responsiveness to role framing, sentence-level stance content density assessed by an LLM judge panel, and cross-condition stance consistency. Evaluating eight frontier models from five vendors, we find that epistemic flexibility is largely orthogonal to general model capability: a 27B open-weight model matches the strongest proprietary systems, the flagship model of one family underperforms its lightweight counterpart, and reasoning-optimized models do not consistently exhibit higher flexibility. Stance content density provides the strongest signal, while surface-level lexical markers such as'I think'can change substantially without corresponding changes in expressed stance. We provide item-level bootstrap confidence intervals, weight-sensitivity analyses, and an explicit discussion of the interpretation limits of the composite score. ESFP measures a model's propensity to adapt its epistemic stance under changing attribution conditions, rather than a general competence measure.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Do Large Language Models Know What They Don't Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty

This work introduces the Epistemic Honesty Quotient (EHQ), which reports three observable sub-scores across two operational axes (epistemic restraint and substantive-answer calibration), and constructs EHQ-3000, a 3,000-question benchmark spanning Fabricated Entity, Post-Cutoff Event, Hyper-Niche True, and Context-Cond...

Ali Şenol, H. Bernard, Huan Liu · 2 citations
Preprint Aug 2026

Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility

It is found that answers with smaller cross-contextual shifts are more likely to be correct or factual, and C3 provides a complementary axis of evaluation and can serve as a benchmark usefulness diagnostic, identifying which portions of a benchmark remain informative even when aggregated scores are widely considered"sa...

Si-Yang Wu, Yi-Bo Jiang, Bryon Aragam · 0 citations
#natural language process... Preprint Sep 2026

When Models Defer to Wrong Answers: A Robustness Audit of Source-Attributed Cues in Multiple-Choice QA

Language models often receive a question together with a claim about what another source answered. We audit whether such claims destabilize answers in multiple-choice question answering. For each item, we hold one wrong option fixed across misleading conditions and vary the cue template attached to it. We introduce \em...

Manikandan Ravikiran, Siddharth Vohra · 0 citations
Open access 2026

Response Stability of Large Language Models Under Meaning-Preserving Prompt Variation

Single-prompt assessments provide limited evidence about whether task-relevant response properties remain stable when the same communicative intention is reformulated. This study examined response stability under controlled, meaning-preserving English prompt variation using 648 responses from 12 researcher-selected arg...

Idrees, Yong-Zhi Liu · 0 citations
Preprint Aug 2026

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.

V. Rodionov, Shamil Assylbekov · 0 citations
#artificial intelligence Preprint Sep 2026

StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions

A dual-framing protocol with minimally varied prompts that use either support- or elimination-oriented framing while keeping the evaluation target fixed is introduced, providing intervention-based evidence that the activations are behaviorally relevant.

Chao Gao, Haijiang Liu, Qiyuan Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.