This work shows that truth representations of a proposition are significantly swayed by partner assertions about that proposition, even when the LLM has enough evidence to determine its truth, and finds evidence that propositions near the decision boundary are more susceptible to having their truth shifted through partner assertions.
Abstract
Prior work has shown that LLMs encode the truth of factual propositions along linear directions in activation space. It's unclear how these representations extend to contextual truth: propositions whose truth is determined by in-context evidence rather than world knowledge. We show that LLMs maintain a linear representation of contextual truth that persists across structurally different output policies, even when the output doesn't require the model to determine a proposition's truth, and show causal evidence via steering experiments. Using the transcripts from a collaborative vision-language task that requires two LLMs to maintain a shared common ground, we show that truth representations of a proposition are significantly swayed by partner assertions about that proposition, even when the LLM has enough evidence to determine its truth. We find evidence that propositions near the decision boundary are more susceptible to having their truth shifted through partner assertions. Separating representation from output distinguish two forms of sycophancy that output behavior alone cannot: the model may accommodate a false proposition while continuing to represent it as false, or shift its representation across the boundary. The latter is 2.59x more common when the model agrees by restating the false claim explicitly than when it agrees implicitly.
It is shown how fact-checking, a generally desirable behavior, can interfere with belief tracking in LLMs and how suppressing this attention at decoding time recovers accuracy only partially and only in some models, calling for future work on intervention methods.
Quang Minh Nguyen, Luis Frentzen Salim· 0 citations
Twin Worlds (TW), a framework for improving reliability in knowledge-intensive reasoning through equivariance-based abstention, is proposed, which identifies when answers are not reliably grounded in the provided evidence and outperforms uncertainty- and sufficiency-based baselines.
Vy Nguyen, Zi-Qi Xu, Jeffrey Chan et al.· 1 citation
It is found that while all models show sensitivity to existential presupposition across syntactic embeddings, determiner types and contextual cues, their behaviour differs markedly in strength and systematicity, with NLI-fine-tuned autoregressive models exhibiting the most coherent and stable projection patterns.
Marie-Léontine Wörgötter, Shiyang Lai, Sebastian Schuster· International Conference on...· 0 citations
Language can describe states of affairs that are false and states of affairs that could not be the case at all. Whether an AI model internally distinguishes these failures remains unclear. I report an exploratory activation study of the multimodal open-weight model Gemma 3 4B IT using 85 prompts from 17 philosophical f...
SymboUQ is a symbolic uncertainty quantification framework that estimates final-answer reliability from reasoning traces by distinguishing symbolizability, whether a claim can be represented in the verifier's formal language, from semantic determinacy, whether its execution yields an entailed or contradicted verdict ra...
Da-Hai Yu, Lin Jiang, Rong-Chao Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.