Coherentist Probabilistic Compositionalism (CPC), an interpretive framework that grounds transformer computation in coherentist theories of interpretation and describes it through four operator roles, is introduced.
Abstract
Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks and architectures. We introduce Coherentist Probabilistic Compositionalism (CPC), an interpretive framework that grounds transformer computation in coherentist theories of interpretation and describes it through four operator roles. Alignment identifies candidate relations, unification integrates supporting information, suppression reduces incompatible alternatives, and routing carries selected information to the output. Across 15 models from five architecture families, the suppression, unification, and routing weight-space signatures correlate with held-out activation-level role measures above random baselines. Suppression is more stable across tasks than unification. Ablating alignment heads reduces downstream suppressive activity beyond a random-head control in 10 models, but similar effects on no-conflict prompts indicate a general upstream dependency, not contradiction-specific coupling. Explicit contradictions significantly shift a layerwise coherence proxy in 14 models; after removing shared residual covariance, the gap has the predicted direction in every model. Base and instruction-tuned variants preserve induction-head score structure ($r{\geq}0.98$) without a consistent shift of operator signatures towards later layers. These results support CPC as a shared vocabulary for comparing transformer mechanisms while showing that their depth and geometric expression remain architecture-specific.
Transformer-based language models perform well on symbolic tasks, yet it remains unclear whether they learn generalizable rules or rely on statistical shortcuts. Mechanistic studies link algorithmic behavior to structured internal representations, motivating the hypothesis that robust reasoning benefits from separating...
Cristina V. Lopes, Yuan-Gang Li, Justin Tian Jin Chen et al.· 0 citations
Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds...
Jia-Le Dai, Hong-Can Deng, Liu-Xian Ma et al.· 0 citations
This work studies activations from ensembles of prompts that require inferring relationships between three tokens corresponding to a cyclic concept to correctly predict the next token and finds a consistent layerwise progression in how the joint dependence among the tokens is geometrically organized and causally used.
Gurbir Arora, To-Ni J. B. Liu, Jiajun Bao et al.· 0 citations
Synthetic semantic-preserving transformations that are rule-based and invertible are used as a probe of alignment generalization and suggest that semantic-preserving distribution shifts can expose a recurring gap in how utility and alignment generalize in current LLMs.
Mo-Han Li, Cheng-Yu Yu, Francesco Sovrano et al.· 0 citations
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal decomposition, and deduction. Although these operations are explicitly distinguished in text, little is known about how they are geometrically organized in representation spaces. To this end, we investiga...
Seogyeong Jeong, Jaehui Hwang, Dongyoon Han et al.· 0 citations
Large Language Models (LLMs) demonstrate remarkable problem-solving capabilities when guided by Chain-of-Thought (CoT) prompting, yet the internal mechanisms underlying these improvements remain poorly understood. In this work, we investigate where CoT-related causal effects emerge across the generated reasoning trajec...
Murat Dura, Serkan Öztürk, Selma Tekir· 1 citation· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.